Ops Journal · Practical notes on running software in production

A backup you have never restored is not a backup

Published 2026-09-24 · 9 min read

The uncomfortable truth about backups is that the backup job is not the thing that has to work. The restore is. A nightly tar that has run cleanly for two years tells you nothing about whether the archive can be turned back into a working system, because the failure modes that matter — a truncated file, a missing exclude, a database dumped while it was being written — all stay invisible until you try to read the data back.

This is the drill that finds those failures on a quiet afternoon instead of during an incident.

Define what a successful restore means before you test

"We restored the backup" is not a criterion, because it does not say what working looks like. Write it down as a checkable statement first. For a typical web service, that is something like:

  • the application starts from the restored files
  • the database answers a known query and returns the expected row count
  • a file uploaded after the last full backup is present

Without that list, a drill degenerates into "the tar file extracted without error", which is the weakest possible signal.

Verify the archive before trusting its contents

For a tar archive, listing the members is cheap and catches truncation immediately:

tar -tzf /backups/srv-2026-09-23.tar.gz > /dev/null && echo "archive readable"

A truncated gzip fails here with unexpected end of file, which is exactly the class of corruption that a successful tar -czf run does not reveal — the producing process can exit zero after writing a short file if the disk filled mid-write.

For anything where integrity matters, keep a checksum alongside the archive and verify it as part of the drill:

sha256sum /backups/srv-2026-09-23.tar.gz > /backups/srv-2026-09-23.sha256

# later, and this is the step that actually proves the bytes survived
sha256sum -c /backups/srv-2026-09-23.sha256

sha256sum -c returns non-zero and prints FAILED on mismatch, which makes it usable directly in a script.

Restore into a scratch location, never over the live system

The drill must not be able to damage production. Extract into a temporary tree and work from there:

sudo mkdir -p /var/tmp/restore-drill
sudo tar -xzf /backups/srv-2026-09-23.tar.gz -C /var/tmp/restore-drill

Then run the checks from your list against that tree. For a config-driven service, the strongest cheap check is to validate the restored configuration rather than start it:

# example: validate an nginx config from the restored tree
nginx -t -c /var/tmp/restore-drill/etc/nginx/nginx.conf

That catches the failure mode where the backup captured files but not the symlinks or permissions they depend on, which is common when the backup was taken with a tool that follows links or drops ownership.

Databases need their own procedure

A file-level copy of a running database is not a backup. It is a snapshot of whatever pages happened to be on disk at that moment, and it is frequently unrecoverable. Use the engine's own dump:

# PostgreSQL: logical dump, consistent without stopping the server
pg_dump -Fc -f /backups/app-2026-09-23.dump appdb

# verify it can be read back, without importing it anywhere
pg_restore --list /backups/app-2026-09-23.dump | head

pg_restore --list reads the archive's table of contents. If the dump is truncated it fails here, and it does so without needing a database to restore into, which makes it a good daily automated check.

The corresponding restore into a scratch database, which is what the drill should exercise at least monthly:

createdb appdb_restore_test
pg_restore -d appdb_restore_test /backups/app-2026-09-23.dump
psql -d appdb_restore_test -c 'select count(*) from users;'
dropdb appdb_restore_test

The row count is the point. It is a number you can compare against production, which turns "the restore ran" into "the data is there".

Automate the cheap checks, do the full drill by hand

A useful split, based on how often each check should run:

Check Frequency Cost
Archive readable (tar -tzf) every backup seconds
Checksum verified every backup minutes on large archives
Dump table of contents readable every backup seconds
Full restore into scratch + row count monthly minutes
Restore on a clean machine quarterly hours

The first three belong in the backup script itself, so a bad backup is caught the same night it is produced rather than a month later. The last two are exercises that prove the procedure works end to end, and they are the ones that find the missing step in your runbook.

Record the time it takes

The number nobody has until they need it is how long a restore actually takes. Time the drill:

time sudo tar -xzf /backups/srv-2026-09-23.tar.gz -C /var/tmp/restore-drill

That figure becomes your recovery time estimate, and it is usually larger than people assume — extraction of a multi-gigabyte archive on a busy disk can take longer than the restore itself. Knowing it in advance changes how you respond when the real thing happens.

Clean up, and check the disk first

Drills consume disk. On a small volume, extracting a full backup can fill the filesystem and cause the very incident you are rehearsing for:

df -h /var/tmp
sudo rm -rf /var/tmp/restore-drill

The minimum viable version

If a full drill is too much to start with, do this much today: pick the most recent backup, list the archive, verify the checksum, and extract one file you can identify. That takes five minutes and will already find the backups that are silently empty, which is the single most common failure.

Then put the monthly restore in a calendar with the name of a person attached to it. A drill that is nobody's job does not happen.