Ops Journal · Practical notes on running software in production

Crash loops: reading systemd restart policies instead of guessing

Published 2026-10-07 · 7 min read

A service that keeps restarting is either recovering badly or failing to start, and the difference is in the unit's restart policy. Reading the policy that is actually in effect — rather than the one in the file you edited — is the first step, and it takes one command.

Read the effective values, not the file

systemd merges unit files from several directories and applies drop-ins. The file you are looking at may not be the one being used. Ask systemd instead:

systemctl show myapp.service -p Restart,RestartUSec,StartLimitBurst,StartLimitIntervalUSec,NRestarts

On this machine, for ssh:

systemctl show ssh -p Restart,RestartUSec,StartLimitBurst
Restart=on-failure
RestartUSec=100ms
StartLimitBurst=5

Each of those matters and each has a default that surprises people.

What each setting does

Restart= decides when systemd brings the service back:

Value Restarts on Typical use
no never one-shot jobs
on-failure non-zero exit, signal, or timeout services that may fail transiently
always any exit, including clean services that should never stop
on-abnormal signals and timeouts only avoid restarting on a clean exit

The distinction between on-failure and always is the one that matters in a crash loop. With on-failure, a service that exits cleanly stays down, which is often what you want — it means the service decided it was done. With always, it comes back regardless, which turns an intentional exit into an endless loop.

RestartUSec is the delay before restarting. The default is 100ms, and that is the reason a broken service can generate hundreds of log lines in a few seconds. A crash loop at 100ms looks like a flood; the same loop at RestartSec=10s is readable.

StartLimitBurst and StartLimitIntervalUSec together define the rate limit: if the service restarts more than StartLimitBurst times within the interval, systemd gives up and puts the unit in a failed state. The default burst is 5 within 10 seconds.

Why the service eventually stops trying

This is the behaviour people misread as "systemd gave up for no reason". Once the burst limit is hit, the unit enters failed and systemd stops restarting it — deliberately, so a broken service cannot spin forever consuming resources and filling logs.

Check whether that has happened:

systemctl status myapp.service --no-pager | head -20
systemctl is-failed myapp.service

The status output includes a line naming the limit that was hit, which tells you it is a crash loop rather than a one-off failure.

To reset the counter and allow restarts again, after fixing the underlying problem:

sudo systemctl reset-failed myapp.service
sudo systemctl start myapp.service

Resetting without fixing the cause simply restarts the loop.

Reading the loop to find the cause

The journal holds every attempt with its exit code, and the sequence is more informative than any single entry:

journalctl -u myapp.service -n 50 --no-pager

Look at the exit status on each attempt. Three patterns:

  • Same non-zero code every time — a deterministic startup failure. Config error, missing file, bad credentials. The code names the class.
  • Signal (for example status=11/SEGV) — the process is crashing. Code problem or corrupt input.
  • status=1/FAILURE after a successful start — the service starts, then exits. Often a foreground/background mismatch: the unit expects the process to stay in the foreground and it is daemonising instead.

That third case is common and worth recognising, because the logs show a successful startup immediately followed by a restart, which looks like a crash but is not.

The foreground/background trap

systemd expects the process it starts to remain in the foreground. A program that forks and returns — the classic daemon pattern — makes systemd think the service exited immediately, and with Restart=always it restarts forever.

The fix is in the service, not the unit, in most cases: run it in the foreground with the flag the program provides for it. Where that is not possible, Type= can be set to match the behaviour, but foreground is the correct answer and the others are workarounds.

A policy that degrades gracefully

For a service where a slow restart is acceptable, this combination keeps a crash loop from becoming a flood while still recovering from transient failures:

[Service]
Restart=on-failure
RestartSec=5
# give up after 5 failures in 5 minutes
StartLimitBurst=5
StartLimitIntervalSec=300

The longer interval is the important change from the default. Five failures in ten seconds is a hair trigger; five in five minutes distinguishes a genuinely broken service from one that tripped once.

Note that StartLimitIntervalSec belongs in the [Unit] section on current systemd, not [Service]. Placing it in the wrong section means it is silently ignored, and the default applies — which is why the effective-value check at the top of this article matters.

Confirm the change took effect

After editing, reload and verify the values systemd is using, not the values in the file:

sudo systemctl daemon-reload
systemctl show myapp.service -p Restart,RestartUSec,StartLimitBurst

Then trigger a failure deliberately and watch one cycle, rather than waiting for a real one:

sudo systemctl restart myapp.service
journalctl -u myapp.service -f

Seeing a single clean restart attempt with the delay you configured is the proof that the policy is live.