Ops Journal · Practical notes on running software in production

Testing whether a remote port is reachable, and what each failure means

Published 2026-10-06 · 7 min read

When a service is unreachable, the failure message contains most of the diagnosis. The mistake is treating "it does not connect" as a single problem when it is three distinct ones with different causes and different fixes.

The three failures and what they mean

Symptom Meaning Where to look
Connection refused Something replied and said no Service not listening, or bound to the wrong interface
Connection timed out Nothing replied at all Firewall dropping packets, or wrong address
No route to host The network itself cannot reach it Routing, wrong subnet, host down

A refusal is the most informative and the most often misread. A refused connection means a packet arrived at a host and that host actively rejected it with a TCP RST. The network path is fine, DNS resolved, and something is running on that address. The problem is the service or the bind address.

A timeout means silence. Either the packet was dropped by a firewall, or it went to an address where nothing is listening but where the stack is configured to drop rather than refuse. Timeouts are the hardest to localise because the failure is the absence of a response.

Test the port without assuming what is running there

The most useful tool is one that only opens a TCP connection and reports what happened:

timeout 5 bash -c 'cat < /dev/null > /dev/tcp/10.0.0.5/5432' && echo "OPEN" || echo "CLOSED or FILTERED"

The /dev/tcp pseudo-device is a bash built-in; it does not exist in sh or dash, which is why the command is wrapped in bash -c. If the connection opens, the redirect succeeds and the command prints OPEN. If it is refused, bash reports the error and the || branch runs.

To distinguish refused from filtered, capture the actual error rather than collapsing it:

timeout 5 bash -c 'cat < /dev/null > /dev/tcp/10.0.0.5/5432'
echo "exit=$?"

The error text is the diagnosis: Connection refused points at the service, Connection timed out points at a firewall, and No route to host points at routing.

Prefer nc or ss when they are available

If netcat is installed it gives the same information with clearer flags:

nc -zv 10.0.0.5 5432          # -z scan only, -v verbose
nc -zvw3 10.0.0.5 5432        # with a 3 second timeout

-z is what makes it a port test rather than a data exchange — without it, nc opens a connection and waits for input, which looks like a hang.

ss can test from the local machine's perspective and show the resulting socket state, which is useful for confirming a connection is being attempted at all:

ss -tan | grep 10.0.0.5

A line in SYN-SENT that persists means the SYN went out and nothing came back, which is the signature of a silent firewall drop.

Localise: is it the service or the path

The single most useful comparison is between connecting locally and connecting remotely.

From the server itself:

ss -tlnp | grep 5432

If the service is listening on 127.0.0.1:5432, it is reachable only from the machine itself and every remote attempt will time out or be refused regardless of firewall rules. This is the most common cause of "the database is down" when the database is in fact running perfectly and bound to loopback.

If it listens on 0.0.0.0:5432 or *:5432 but remote connections fail, the problem is between the hosts. Check the host firewall:

sudo iptables -L -n -v 2>/dev/null | head -20
sudo nft list ruleset 2>/dev/null | head -30

Then check whether a cloud provider's security group or network ACL is filtering, which no amount of local inspection will reveal.

DNS is a separate question

A timeout caused by a wrong address looks identical to a timeout caused by a firewall. Resolve the name first and confirm it is what you expect:

getent hosts db.internal

getent hosts is preferable to nslookup for this because it uses the same resolution path as the application, including /etc/hosts and NSS configuration. An application resolving a name differently from your diagnostic tool is a real and confusing class of bug, and getent avoids it.

To compare the address actually being used, connect by IP as well as by name:

nc -zvw3 10.0.0.5 5432     # by address
nc -zvw3 db.internal 5432  # by name

If the IP works and the name does not, the problem is DNS or /etc/hosts, not the service.

What to record when it works

The value of this exercise is having a baseline. When the service is healthy, record what a successful connection looks like:

timeout 5 bash -c 'cat < /dev/null > /dev/tcp/db.internal/5432' && echo "reachable $(date -u +%FT%TZ)"

A log line like that, captured on a schedule, turns "it was working yesterday" into a timestamp you can correlate against the change that broke it. That correlation is usually the whole investigation.