Ops Journal · Practical notes on running software in production

Too many open files: finding a file descriptor leak before it takes the service down

Published 2026-10-07 · 8 min read

Too many open files is one of the few errors that tells you exactly what is wrong and nothing about where. The limit was reached, a service cannot open anything else, and from the outside it looks like a sudden failure even though the cause accumulated over days.

The work is in three parts: find the limit, find which process is at it, and find what it is holding.

Find the limit, and notice there are two

A process has a soft limit and a hard limit. The soft limit is what is enforced and can be raised by the process itself up to the hard limit. The hard limit is the ceiling, and only root can raise it.

ulimit -n        # soft limit
ulimit -Hn       # hard limit

On this machine, in a normal shell:

4096
1048576

That gap is the usual situation on a modern systemd host: the shell's soft limit is conservative while the hard limit is effectively unlimited. A service that inherits the soft limit will hit a wall at 4096 while the system is nowhere near its own capacity.

For a process already running, ask the kernel rather than the shell — the limits are per-process and may have been set at start time:

cat /proc/<PID>/limits | grep -i 'open files'

This reads the actual enforced values for that process, which is the only authoritative answer. A service started by systemd may have a different limit from one started by hand, because systemd applies LimitNOFILE from the unit.

Find which process is closest to its limit

Rather than guessing, count descriptors for the processes you care about. The count is a directory listing of the process's own descriptor table:

for pid in $(pgrep -f 'myapp'); do
  n=$(sudo ls /proc/$pid/fd 2>/dev/null | wc -l)
  lim=$(grep -i 'open files' /proc/$pid/limits | awk '{print $4}')
  printf 'pid=%-8s open=%-8s soft_limit=%s
' "$pid" "$n" "$lim"
done

The comparison between open and soft_limit is the whole diagnosis. A service sitting at 3,900 of 4,096 has minutes left; one at 40 is not the problem.

What a healthy process looks like

To judge whether a number is a leak, you need a baseline. For a service that has been running for days, compare the count now with the count shortly after startup. A stable service fluctuates within a range; a leaking one climbs monotonically.

Sampling it over an hour is enough to see the trend:

pid=$(pgrep -f 'myapp' | head -1)
for i in $(seq 1 6); do
  printf '%s open=%s
' "$(date -u +%H:%M)" "$(sudo ls /proc/$pid/fd | wc -l)"
  sleep 600
done

A count that only ever goes up — even slowly — is a leak. A count that rises under load and falls when idle is normal connection churn.

Find what is being held

Listing the descriptors shows what each one points at, which usually names the culprit directly:

sudo ls -l /proc/<PID>/fd | head -30

Each line shows the target: a file path, a socket, a pipe, or [eventpoll]. Reading the breakdown by type is faster than reading individual lines:

sudo ls -l /proc/<PID>/fd \
  | awk '{print $NF}' \
  | sed -E 's#.*\.(log|db|sock)$#\1#; s#socket:.*#socket#; s#pipe:.*#pipe#; s#anon_inode.*#anon_inode#' \
  | sort | uniq -c | sort -rn | head

The two leaks worth knowing:

  • Sockets climbing without bound — almost always a connection pool that never closes connections on error paths. The application opens a new connection on failure and forgets the old one.
  • The same file opened many times — a file handle opened per request and never closed, often a config or log file re-opened inside a hot loop.

The kernel-side view, when a process will not show its files

If /proc/<PID>/fd is unreadable (a process in D state, or one owned by another user without sudo), lsof gives the same information from the kernel's file table:

sudo lsof -p <PID> | wc -l
sudo lsof -p <PID> | awk '{print $5}' | sort | uniq -c | sort -rn | head

lsof is slower than reading /proc directly because it walks the whole table, but it works when the direct read fails and it can be filtered by type.

Fixing it: raise the limit, then fix the leak

Raising LimitNOFILE buys time and is the right immediate action, but it is not a fix — it moves the wall further out while the leak continues.

In a systemd unit:

[Service]
LimitNOFILE=65535

Then:

sudo systemctl daemon-reload
sudo systemctl restart myapp.service

Confirm the new limit actually applied, rather than assuming:

cat /proc/$(pgrep -f myapp | head -1)/limits | grep -i 'open files'

A limit set in a shell profile will not affect a systemd service, and this check is how you find that out before the next incident.

The real fix is on the application side: ensure every error path closes what it opened. Where that is not under your control, a restart on a schedule is a legitimate mitigation, though it is worth being honest that it is a mitigation and not a repair.

Add the alarm before you need it

The useful alert is on the ratio, not on the error. Watching open / soft_limit and warning at 80% gives time to act, while alerting on EMFILE means the service has already failed:

# exit 1 when a process is above 80% of its descriptor limit
pid=$(pgrep -f myapp | head -1)
open=$(sudo ls /proc/$pid/fd | wc -l)
soft=$(grep -i 'open files' /proc/$pid/limits | awk '{print $4}')
[ "$open" -gt $((soft * 80 / 100)) ] && echo "WARN $open/$soft"

That is the same three numbers as the diagnosis, arranged so they run unattended.