Reading a disk-full incident in the right order
A full disk is a race: the machine may be minutes from refusing writes, and the temptation is to start deleting things immediately. Doing the diagnosis in a fixed order is faster than improvising, because each step either finds the cause or eliminates a whole class of them.
Step one: is it space or inodes
These are different resources and produce nearly identical symptoms. A filesystem can be completely out of inodes with terabytes free, and the error messages are similar enough to mislead.
df -h /
df -i /
Filesystem Size Used Avail Use% Mounted on
/dev/vda1 232G 70G 162G 31% /
Filesystem Inodes IUsed IFree IUse% Mounted on
/dev/vda1 31326208 1476771 29849437 5% /
The IUse% column on the second command is the one people forget to check. A
filesystem at 100% inodes with 31% space used fails writes with the same
No space left on device, and no amount of freeing space fixes it — the answer
is finding the directory holding millions of tiny files.
A directory that accumulates small files is almost always a session store, a cache, or a mail queue. Counting entries per directory finds it:
sudo find /var -xdev -type d -printf '%p
' 2>/dev/null | head
For a targeted hunt, count files in the usual suspects:
for d in /var/spool /var/lib /tmp /var/tmp; do
printf '%-16s %s
' "$d" "$(sudo find $d -xdev -type f 2>/dev/null | wc -l)"
done
Step two: which filesystem is actually full
On a machine with several mounts, / being fine says nothing about the mount
holding the data:
df -h --output=source,size,used,avail,pcent,target | sort -k5 -r
Sorting by the use-percent column puts the worst first. This matters because the
obvious reaction — cleaning up /var/log — does nothing when the full filesystem
is /srv or a mounted volume.
Step three: where inside that filesystem
Now, and only now, walk the directory that is full. One level at a time is faster than a deep recursive scan and almost always enough:
du -xhd1 /var | sort -rh | head -6
3.6G /var
3.0G /var/log
352M /var/cache
254M /var/lib
2.2M /var/backups
28K /var/tmp
The -x flag keeps the walk on one filesystem. Without it, du descends into
other mounts and reports sizes that belong to disks that are not full, which
produces the confusing result where the numbers add up to far more than the
filesystem's total size.
Descend into the winner and repeat:
du -xhd1 /var/log | sort -rh | head -8
Step four: when du does not explain df
This is the step that resolves the puzzling cases. If df says the filesystem is
full but du cannot account for the space, there are deleted files still held
open by a running process. The directory entry is gone; the blocks are not freed
until the last file descriptor closes.
sudo lsof +L1 2>/dev/null | head -20
+L1 means "list files with a link count below one" — that is, unlinked files
that still exist because something has them open. The SIZE column shows how
much each is holding. A log file deleted during an incident that is still held by
the logging daemon shows up here, and it explains the gap exactly.
The fix is to restart the process holding the file, or signal it to reopen its logs. Deleting more files frees nothing:
# for a service that reopens logs on SIGHUP
sudo systemctl reload myapp.service
Step five: check for a mounted-over directory
The other case where du and df disagree is a directory that has had a
filesystem mounted over it. Files written before the mount still occupy space on
the underlying filesystem but are invisible because the mount hides them.
The check is to compare the mount table against the directory contents:
findmnt -o TARGET,SOURCE,FSTYPE | head -20
If a path that should contain data shows a different source than you expect, or a directory that should be empty on the parent filesystem is holding gigabytes, unmounting will reveal the hidden files. This is worth knowing because it is non-obvious and looks like a filesystem that is lying about its usage.
Step six: reclaim, in order of safety
With the cause identified, free space starting with the things that are guaranteed safe:
# journal, capped rather than deleted wholesale
sudo journalctl --vacuum-size=500M
# package cache: always safe, re-downloadable
sudo apt-get clean
# old rotated logs
sudo find /var/log -type f -name '*.gz' -mtime +30 -delete
Then set the limits that stop it recurring — SystemMaxUse in
/etc/systemd/journald.conf for the journal, a maxsize in the relevant
logrotate rule for text logs, and a disk-space alert at 80% rather than 95% so
the next incident is caught while there is still room to work.
The order is the point
df for space or inodes, df again for which mount, du for where inside it,
lsof +L1 when they disagree, then reclaim. Skipping to the last step is what
turns a five-minute diagnosis into an hour of deleting the wrong files.