Ops Journal · Practical notes on running software in production

A health check that does not lie: checking the thing that can actually fail

Published 2026-10-03 · 7 min read

The purpose of a health check is to answer one question: is this service able to do its job right now? Most health checks answer a different and easier question — is the process listening — and the gap between those two is where outages hide.

A service can be listening, returning 200 on /, and completely unable to serve a request that touches its database. That is the failure a health check should catch and most do not.

Why the process check is not enough

The failure modes that matter are almost never "the process died". They are:

  • the database connection pool is exhausted
  • a dependency is reachable but returning errors
  • the application is up but its migration did not complete
  • the disk holding its data is full, so writes fail while reads succeed

Every one of those leaves the process running and the port open. A check that only tests the socket passes in all four cases.

Test the dependency, not the socket

For an HTTP service, the first improvement is to follow redirects and treat any non-2xx as failure, which curl does not do by default:

curl -fsS -o /dev/null -m 5 http://127.0.0.1:8080/healthz

The flags are worth knowing individually: -f makes curl exit non-zero on an HTTP error instead of printing the error page and exiting zero, -s silences the progress meter, -S still shows a real error when one occurs, and -m 5 caps the whole request at five seconds so a hung service fails the check rather than hanging the monitor.

The bigger improvement is to make /healthz do real work. A health endpoint that only returns {"status":"ok"} is decoration. One that queries the database is a check:

# the application's own health endpoint should touch its dependencies;
# verify it does by stopping the dependency and watching it fail
sudo systemctl stop postgresql
curl -fsS -o /dev/null -w '%{http_code}\n' http://127.0.0.1:8080/healthz
sudo systemctl start postgresql

If that still prints 200 with the database stopped, the endpoint is not checking anything and the monitor built on it is blind.

Check the dependency directly when you can

Where the application does not expose a dependency-aware endpoint, a check can test the dependency itself. For a database, the strongest cheap test is a real query, not a TCP connect:

# PostgreSQL: a connection that also proves it can read
pg_isready -h 127.0.0.1 -p 5432 -t 3

# and a query, which catches a server that accepts connections but cannot serve
psql -h 127.0.0.1 -U app -d appdb -tAc 'select 1' >/dev/null

pg_isready answers "is the server accepting connections", which is weaker than a query but costs almost nothing. The select 1 is the stronger check; it fails if the server is up but the database is in recovery or the disk is full.

For a cache, the equivalent is a write and read back, since a cache that accepts commands but cannot evict or persist is still broken:

redis-cli -h 127.0.0.1 ping | grep -q PONG

Make the check fail loudly and stop

A health check that returns non-zero but is run by something that ignores the exit code is the same as no check. If the check runs from systemd, wire it to a failure action so a persistent failure is visible without a dashboard:

# /etc/systemd/system/healthcheck.service
[Unit]
Description=Service health check

[Service]
Type=oneshot
ExecStart=/usr/local/bin/healthcheck.sh

[Install]
WantedBy=multi-user.target

with a timer that runs it every few minutes:

# /etc/systemd/system/healthcheck.timer
[Unit]
Description=Run health check every 5 minutes

[Timer]
OnCalendar=*:0/5
Persistent=true

[Install]
WantedBy=timers.target

The OnCalendar=*:0/5 expression means every five minutes; confirm it with systemd-analyze calendar '*:0/5' before trusting it.

A failed oneshot unit shows up in systemctl --failed and in the journal, which is enough to catch a service that has been unhealthy for hours.

Avoid the check that flaps

Two habits keep a health check from becoming noise. First, require consecutive failures before alerting — a single timeout during a deploy is not an incident. Second, give the check a timeout shorter than its interval, so a hung check cannot overlap with the next run:

# 5 second cap on a check that runs every 5 minutes
timeout 5 /usr/local/bin/healthcheck.sh

timeout from coreutils returns 124 when it has to kill the command, which is a distinct exit code you can alert on separately from a genuine failure.

Test the check by breaking the thing

A health check is code, and code that has never failed is code you do not know works. Break the dependency deliberately, in a controlled way, and confirm the check goes red:

sudo systemctl stop postgresql
/usr/local/bin/healthcheck.sh; echo "exit=$?"
sudo systemctl start postgresql
/usr/local/bin/healthcheck.sh; echo "exit=$?"

The exit codes should differ. If they do not, the check is not testing what you think, and it is worth fixing before it is needed.

What to leave out

A health check should be fast and side-effect free. Avoid checks that write to the database, send email, or call a third-party API — those turn a monitoring system into a source of load and, worse, of failures. If the check itself can break the service, it is not a check.