When /var/log eats the disk: finding what actually grew
A server that has run for a year without attention will eventually page you with
No space left on device. The instinct is to look for a runaway application
writing data files. In practice, on a machine that only runs services and their
logs, the log directory is the answer far more often than not.
Start with the top-level split
Resist the urge to run a recursive du over the whole filesystem. It walks
every file and on a busy box takes minutes, while the answer is usually visible
one level down:
du -xhd1 /var | sort -rh | head -6
On the machine this was written on:
3.6G /var
3.0G /var/log
352M /var/cache
254M /var/lib
2.2M /var/backups
-x keeps the walk on one filesystem, which matters when /var/lib/docker or a
database directory is a separate mount — without it, du descends into those and
reports sizes that have nothing to do with the disk that is actually full. -d1
limits depth to one level, and sort -rh puts the largest first.
/var/log at 3.0G out of 3.6G is the whole story. Now find out which writer owns
it.
Separate the journal from plain log files
Since systemd took over logging, /var/log holds two different things that need
different fixes: the binary journal under /var/log/journal, and traditional
text files written by rsyslog, nginx, or application daemons.
du -xhd1 /var/log | sort -rh | head -8
The journal's own accounting is more reliable than reading the directory size, because the journal can be capped and will report the cap:
journalctl --disk-usage
If you are not in the adm or systemd-journal group you will get a notice
saying you cannot see other users' messages, and the number reported may be
smaller than reality. Run it with sudo when you want the real figure:
sudo journalctl --disk-usage
The journal is the usual suspect, and it has a limit you have not set
By default the journal grows until it hits a percentage of the filesystem
(SystemMaxUse defaults to 10% of the filesystem size, capped at 4G). On a
232G disk that ceiling is 4G, which is enough to be a problem and not enough to
be obvious.
Check the current cap and what is actually in use:
sudo journalctl --disk-usage
grep -E '^\s*SystemMaxUse|^\s*MaxRetentionSec' /etc/systemd/journald.conf
If both lines are commented out, the defaults apply and nothing is holding the
journal down. Set an explicit ceiling in /etc/systemd/journald.conf:
[Journal]
SystemMaxUse=500M
SystemKeepFree=1G
MaxRetentionSec=1month
SystemKeepFree is the more important of the two on a small disk: it tells the
journal to leave at least that much space free regardless of SystemMaxUse. Then
apply it without waiting for a restart:
sudo systemctl restart systemd-journald
Note that restarting systemd-journald does not delete anything. The cap is
enforced when new entries arrive, so the journal shrinks as it rotates rather
than immediately. To reclaim space now, vacuum explicitly:
# keep the last two weeks
sudo journalctl --vacuum-time=2weeks
# or cap the total size
sudo journalctl --vacuum-size=500M
--vacuum-time and --vacuum-size both delete archived journal files. They do
not touch the currently active file, so you may free less than you expect in one
pass; run it twice if the number barely moved.
Text logs need logrotate, and logrotate needs to be running
Text files under /var/log are rotated by logrotate, which on modern Ubuntu is
driven by a systemd timer rather than a cron job. Confirm it is actually firing:
systemctl list-timers logrotate.timer
If the LAST column is days old, rotation has stopped and every file will grow
without bound. The two causes worth checking first are a stale .dpkg-old
suffix on a config file, and a rotation rule whose postrotate script exits
non-zero and makes logrotate skip the remaining configs:
sudo logrotate --debug /etc/logrotate.conf 2>&1 | grep -iE 'error|skipping'
--debug prints what it would do and changes nothing, which is the right way
to test a config before trusting it.
For an application writing its own file with no rotation rule, the minimal fix is
a drop-in in /etc/logrotate.d/:
/var/log/myapp/*.log {
daily
rotate 14
compress
delaycompress
missingok
notifempty
copytruncate
}
copytruncate is the one line that matters for a process you cannot restart: it
copies the file and then truncates it in place, so the daemon keeps its open file
descriptor instead of writing into a file that was renamed out from under it.
A note on what not to do
Deleting files from /var/log by hand while a process holds them open frees no
space. The directory entry disappears, but the blocks stay allocated until the
last file descriptor closes. df and du will disagree, which is the classic
symptom:
# df says the disk is full, du says it is not
df -h / && du -xsh /var/log
If df reports far more usage than du can account for, look for deleted files
still held open:
sudo lsof +L1 2>/dev/null | head -20
The fix is to restart the process holding the file, not to delete more files.
The three commands worth remembering
du -xhd1 /var | sort -rh | head # where did it go
sudo journalctl --disk-usage # is the journal the reason
systemctl list-timers logrotate.timer # is rotation still alive
A disk that filled once will fill again. Setting SystemMaxUse and confirming
the logrotate timer is live takes a few minutes and turns a recurring incident
into a non-event.