A health check that does not lie: checking the thing that can actually fail
The purpose of a health check is to answer one question: is this service able to do its job right now? Most health checks answer a different and easier question — is the process listening — and the gap between those two is where outages hide.
A service can be listening, returning 200 on /, and completely unable to serve
a request that touches its database. That is the failure a health check should
catch and most do not.
Why the process check is not enough
The failure modes that matter are almost never "the process died". They are:
- the database connection pool is exhausted
- a dependency is reachable but returning errors
- the application is up but its migration did not complete
- the disk holding its data is full, so writes fail while reads succeed
Every one of those leaves the process running and the port open. A check that only tests the socket passes in all four cases.
Test the dependency, not the socket
For an HTTP service, the first improvement is to follow redirects and treat any
non-2xx as failure, which curl does not do by default:
curl -fsS -o /dev/null -m 5 http://127.0.0.1:8080/healthz
The flags are worth knowing individually: -f makes curl exit non-zero on an
HTTP error instead of printing the error page and exiting zero, -s silences
the progress meter, -S still shows a real error when one occurs, and -m 5
caps the whole request at five seconds so a hung service fails the check rather
than hanging the monitor.
The bigger improvement is to make /healthz do real work. A health endpoint
that only returns {"status":"ok"} is decoration. One that queries the database
is a check:
# the application's own health endpoint should touch its dependencies;
# verify it does by stopping the dependency and watching it fail
sudo systemctl stop postgresql
curl -fsS -o /dev/null -w '%{http_code}\n' http://127.0.0.1:8080/healthz
sudo systemctl start postgresql
If that still prints 200 with the database stopped, the endpoint is not
checking anything and the monitor built on it is blind.
Check the dependency directly when you can
Where the application does not expose a dependency-aware endpoint, a check can test the dependency itself. For a database, the strongest cheap test is a real query, not a TCP connect:
# PostgreSQL: a connection that also proves it can read
pg_isready -h 127.0.0.1 -p 5432 -t 3
# and a query, which catches a server that accepts connections but cannot serve
psql -h 127.0.0.1 -U app -d appdb -tAc 'select 1' >/dev/null
pg_isready answers "is the server accepting connections", which is weaker than
a query but costs almost nothing. The select 1 is the stronger check; it fails
if the server is up but the database is in recovery or the disk is full.
For a cache, the equivalent is a write and read back, since a cache that accepts commands but cannot evict or persist is still broken:
redis-cli -h 127.0.0.1 ping | grep -q PONG
Make the check fail loudly and stop
A health check that returns non-zero but is run by something that ignores the exit code is the same as no check. If the check runs from systemd, wire it to a failure action so a persistent failure is visible without a dashboard:
# /etc/systemd/system/healthcheck.service
[Unit]
Description=Service health check
[Service]
Type=oneshot
ExecStart=/usr/local/bin/healthcheck.sh
[Install]
WantedBy=multi-user.target
with a timer that runs it every few minutes:
# /etc/systemd/system/healthcheck.timer
[Unit]
Description=Run health check every 5 minutes
[Timer]
OnCalendar=*:0/5
Persistent=true
[Install]
WantedBy=timers.target
The OnCalendar=*:0/5 expression means every five minutes; confirm it with
systemd-analyze calendar '*:0/5' before trusting it.
A failed oneshot unit shows up in systemctl --failed and in the journal, which
is enough to catch a service that has been unhealthy for hours.
Avoid the check that flaps
Two habits keep a health check from becoming noise. First, require consecutive failures before alerting — a single timeout during a deploy is not an incident. Second, give the check a timeout shorter than its interval, so a hung check cannot overlap with the next run:
# 5 second cap on a check that runs every 5 minutes
timeout 5 /usr/local/bin/healthcheck.sh
timeout from coreutils returns 124 when it has to kill the command, which is a
distinct exit code you can alert on separately from a genuine failure.
Test the check by breaking the thing
A health check is code, and code that has never failed is code you do not know works. Break the dependency deliberately, in a controlled way, and confirm the check goes red:
sudo systemctl stop postgresql
/usr/local/bin/healthcheck.sh; echo "exit=$?"
sudo systemctl start postgresql
/usr/local/bin/healthcheck.sh; echo "exit=$?"
The exit codes should differ. If they do not, the check is not testing what you think, and it is worth fixing before it is needed.
What to leave out
A health check should be fast and side-effect free. Avoid checks that write to the database, send email, or call a third-party API — those turn a monitoring system into a source of load and, worse, of failures. If the check itself can break the service, it is not a check.