Troubleshooting flowchartProcesses, resources, ports, firewallFrom symptom to root cause
Linux
Triage
A server misbehaves and every tool is one command away. This is the order to use them in: pin down the symptom, read the logs, check the service, then test CPU, memory, disk, network, ports and the firewall one at a time until one of them explains what you saw. If none does, trace the process itself.
The flowchartTap a box to jump to its module
Define the symptom
Write one sentence before you run a single command: what is wrong, since when, and for whom.
Most wasted hours start with a vague symptom. "The server is slow" sends you everywhere. "Checkout API p95 went from 200 ms to 4 s at 14:05 for all users, other APIs are fine" tells you where to look and what a fix must change. Get the clock, the scope and the last change first.
| Question | Why it narrows the search | How to answer it |
|---|---|---|
| What exactly fails? | Errors, slowness and timeouts lead to different modules. | Reproduce it: curl -sv -o /dev/null -w "%{http_code} %{time_total}\n" URL |
| Since when? | Lets you line up logs, deploys and cron jobs. | journalctl --since "14:00" --until "14:15" |
| How widespread? | One user, one host, or everyone points at client, node or shared dependency. | Check a second host and a second endpoint. |
| What changed? | Most incidents follow a deploy, config edit or package upgrade. | grep -E "install|upgrade" /var/log/dpkg.log | tail |
| Is it still happening? | A past spike needs history; a live one needs live tools. | uptime, journalctl -f |
Sixty seconds of context
Run these first on any box you just logged into. They tell you how long it has been up, how loaded it is, who else is on it, and whether the kernel has complained.
uptime # uptime and load averages
who # who else is logged in right now
last -n 5 reboot # recent reboots
dmesg -T --level=err,warn | tail -n 20
systemctl list-units --failed$ uptime 14:12:03 up 41 days, 3:10, 2 users, load average: 7.92, 6.10, 2.41 $ systemctl list-units --failed UNIT LOAD ACTIVE SUB DESCRIPTION ● letterpad.service loaded failed failed letterpad.io NestJS API 1 loaded units listed.
Read load averages left to right: the last 1, 5 and 15 minutes. 7.92, 6.10, 2.41 on a 4 core box means it got busy recently and is still climbing. Compare with nproc.
Check logs
Errors from this boot first, then narrow by unit, time and text. The kernel log tells you about OOM kills and disk errors.
journalctl -p err -b # errors and worse since boot
journalctl -u nginx -u letterpad --since "15 min ago"
journalctl -k -b # kernel only (same as dmesg)
journalctl -b -1 -e # end of the previous boot
journalctl -u letterpad --grep "ECONN|timeout" -n 50$ journalctl -p err -b --no-pager | tail -n 4 Oct 10 14:05:11 web-1 kernel: Out of memory: Killed process 2241 (node) total-vm:2843120kB, anon-rss:1530212kB Oct 10 14:05:11 web-1 systemd[1]: letterpad.service: A process of this unit has been killed by the OOM killer. Oct 10 14:05:11 web-1 systemd[1]: letterpad.service: Failed with result 'oom-kill'. Oct 10 14:05:16 web-1 systemd[1]: letterpad.service: Scheduled restart job, restart counter is at 3.
| Flag | Meaning |
|---|---|
-p err | Priority error and worse. Levels: emerg alert crit err warning notice info debug. |
-b, -b -1 | This boot, the previous boot. --list-boots shows them all. |
-u UNIT | One service. Repeat it to interleave several on one timeline. |
-k | Kernel messages: OOM killer, disk I/O errors, segfaults, link up and down. |
--since / --until | A window: "2026-10-10 14:00", "1 hour ago", today. |
-f, -e, -n N | Follow live, jump to the end, last N lines. |
--grep | Regex on the message text. |
-o json-pretty | Every field, including _PID, _EXE and _SYSTEMD_UNIT. |
Files that are not in the journal
| File | What lands there |
|---|---|
/var/log/nginx/*.log | Access and error logs. 502 and 504 explanations live in the error log. |
/var/log/auth.log | SSH logins, sudo use, failed passwords. |
/var/log/syslog | Ubuntu's plain text copy of most of the journal. |
/var/log/postgresql/ | Slow queries, lock waits, connection limit errors. |
/var/log/apt/history.log | What was installed or upgraded, and when. |
Check service state
Is the process running, restarting, stuck or a zombie? systemd tells you about the unit; ps tells you about every process.
systemctl status letterpad
systemctl show letterpad -p ActiveState -p SubState -p NRestarts -p Result
ps -eo pid,ppid,user,stat,%cpu,%mem,etime,cmd --sort=-%cpu | head
pstree -p $(systemctl show -p MainPID --value letterpad)
pgrep -a node$ ps -eo pid,ppid,user,stat,%cpu,%mem,etime,cmd --sort=-%cpu | head -n 4 PID PPID USER STAT %CPU %MEM ELAPSED CMD 2318 1 letterpad Rsl 97.4 18.2 02:11 /usr/bin/node dist/main.js 911 1 postgres Ss 3.1 2.4 41-03:09:55 /usr/lib/postgresql/16/bin/postgres 1312 1 rabbitmq Ssl 1.0 3.8 41-03:09:51 /usr/lib/erlang/erts/bin/beam.smp
Read the STAT column
| State | Meaning | What to do |
|---|---|---|
R | Running or ready to run. | Normal. Many R processes plus high load means CPU pressure. |
S | Sleeping, waiting for an event. | Normal for idle servers. |
D | Uninterruptible sleep, usually waiting on disk or NFS. | Cannot be killed. Look at module 06. Many D processes push load up without CPU use. |
Z | Zombie: exited, parent has not collected it. | Harmless alone. Thousands mean the parent is buggy; restart the parent. |
T | Stopped by a signal or a debugger. | kill -CONT PID resumes it. |
s l + < N | Session leader, multi-threaded, foreground, high and low priority. | Extra flags after the state letter. |
Signals you will actually send
| Command | Signal | Effect |
|---|---|---|
kill PID | TERM (15) | Ask to exit cleanly. Always try first. |
kill -HUP PID | HUP (1) | Many daemons reload config on it (nginx, sshd). |
kill -INT PID | INT (2) | Same as Ctrl+C. |
kill -9 PID | KILL (9) | Immediate, no cleanup. Last resort, and useless on D state. |
kill -STOP / -CONT | STOP, CONT | Pause and resume. |
pkill -f "dist/main.js" | TERM | Match by full command line. Check with pgrep -af first. |
For a service under systemd, prefer systemctl restart over kill: systemd knows the whole process tree and the restart policy.
Resources: CPU
Check resources one at a time and stop at the first one that is saturated. Start with CPU: who is using it, and is it really user code?
Swipe sideways to see the whole diagram
- 1%usr near 90 means your own code is busy. Find the hot path with a profiler, module 12.
- 2A high %iowait would mean the CPU sits idle waiting on the disk: go to module 06 instead.
- 3
pidstatnames the process: herenode, PID 2318, at 95% of one core.
For each resource ask three things, the USE method: Utilisation (how busy), Saturation (is work queuing) and Errors. Install sysstat once to get pidstat, mpstat and iostat.
sudo apt install -y sysstat htop
nproc # number of CPUs, to compare with load
top -o %CPU # press 1 for per-CPU, P sort by CPU, H threads
mpstat -P ALL 1 3 # per-CPU breakdown, 3 samples
pidstat -u 1 5 # per-process CPU every second
pidstat -t -p 2318 1 # per-thread for one PID$ mpstat 1 1 CPU %usr %nice %sys %iowait %irq %soft %steal %idle all 88.21 0.00 6.02 0.25 0.00 0.75 0.00 4.77 $ pidstat -u 1 1 UID PID %usr %system %CPU CPU Command 998 2318 92.00 3.00 95.00 2 node
| Column | High value means |
|---|---|
%usr | Your code is busy: a hot loop, JSON work, regex, crypto. Profile it (module 12). |
%sys | The kernel is busy for you: many syscalls, context switches, small writes. |
%iowait | CPU idle while waiting for disk. It is a disk problem, go to module 06. |
%steal | The hypervisor gave your CPU to another VM. Nothing to fix inside; resize or move. |
%soft | Packet processing. Very high under heavy network load. |
load average > nproc | Work is queuing. Includes D state processes, so check iowait before blaming CPU. |
Node.js runs your JavaScript on one thread, so one node at 100% on a 4 core box is saturated even though top shows 25% overall. Watch per-process numbers, not just the total.
Resources: memory
Look at available, not free. Then check swap activity and whether the OOM killer has been busy.
Swipe sideways to see the whole diagram
- 1Read available, not free. Linux fills spare RAM with cache and hands it back when asked.
- 2
siandsoabove zero for seconds means the box is swapping right now, and everything slows down. - 3For a Node service, cap the heap below the unit's
MemoryMaxwith--max-old-space-size, so you get a heap error you can read instead of the OOM killer.
free -h
vmstat 1 5 # si/so = swap in/out per second
ps -eo pid,user,rss,vsz,comm --sort=-rss | head
journalctl -k --grep "Out of memory|oom-kill" --since today
cat /proc/2318/status | grep -E "VmRSS|VmSwap|Threads"
systemctl show letterpad -p MemoryCurrent -p MemoryMax$ free -h total used free shared buff/cache available Mem: 3.8Gi 3.5Gi 112Mi 24Mi 240Mi 161Mi Swap: 2.0Gi 1.9Gi 100Mi $ vmstat 1 2 procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu----- r b swpd free buff cache si so bi bo in cs us sy id wa st 3 4 1992344 114820 8204 237600 1840 2210 9120 2380 3110 5902 21 14 9 56 0
| Signal | What it tells you |
|---|---|
available near zero | Real memory pressure. free alone is misleading: Linux uses spare RAM as cache on purpose. |
si / so above zero for seconds | The box is swapping right now. Everything gets slow and wa rises. |
b column above zero | Processes blocked on I/O, often caused by swapping. |
Killed process in the kernel log | The OOM killer picked a victim. The log line names it and its RSS. |
| RSS of one process keeps growing | A leak. Compare ps a few minutes apart, or watch MemoryCurrent. |
For a Node service, cap the heap below the unit's MemoryMax with NODE_OPTIONS=--max-old-space-size=512, so Node throws a heap error you can read instead of the kernel killing it silently.
Resources: disk
Two different problems hide here: a disk that is full (space or inodes) and a disk that is slow (latency and queueing).
Swipe sideways to see the whole diagram
- 1Full:
dfsays 100%. A log removed withrmwhilenodestill has it open keeps its 8.5G until that process lets go. - 2Free it by restarting the process, or truncate it in place with
: > /proc/2318/fd/21. - 3Slow:
aqu-szabove 1 for long means requests are waiting. On NVMe and RAID, trustawaitover%util.
Is it full?
df -hT -x tmpfs -x devtmpfs # space per filesystem
df -i # inodes: full even with free space
sudo du -xh --max-depth=1 /var | sort -h | tail
sudo lsof +L1 # deleted files still held open$ df -hT -x tmpfs -x devtmpfs Filesystem Type Size Used Avail Use% Mounted on /dev/vda1 ext4 48G 48G 0 100% / $ sudo lsof +L1 COMMAND PID USER FD TYPE DEVICE SIZE/OFF NLINK NODE NAME node 2318 letterpad 21w REG 252,1 9126805504 0 51234 /var/lib/letterpad/debug.log (deleted)
lsof +L1 catches the classic trap: someone deleted a huge log with rm, but the process still has it open, so the space is not freed. Restart that process, or truncate it in place with : > /proc/2318/fd/21.
Is it slow?
iostat -xz 1 3 # extended stats, skip idle devices
pidstat -d 1 # read/write per process
sudo iotop -oPa # live, only processes doing I/O$ iostat -xz 1 1 Device r/s w/s rkB/s wkB/s r_await w_await aqu-sz %util vda 412.0 880.0 26368.0 70400.0 18.40 41.90 48.21 99.60
| Column | Read it as |
|---|---|
r_await / w_await | Average ms per request, queue time included. SSDs should be single digit. |
aqu-sz | Average queue length. Above 1 for long means requests are waiting. |
%util | Time the device was busy. Near 100 on a single disk means saturated; on NVMe and RAID it can mislead, trust await. |
Resources: network
Work outward: is the interface up, is there a route, does DNS resolve, does the path drop packets, and where does a request spend its time?
Swipe sideways to see the whole diagram
- 1The
curl -wtimes count from the start, so subtract each from the next to get one phase. - 2A slow
time_namelookuppoints at DNS; a slowtime_connectat distance, packet loss or a full SYN queue. - 3Here DNS, TCP and TLS take 89 ms together. The other 2.8 s is the application or the database.
ip -br addr # interfaces and addresses
ip route get 1.1.1.1 # which route and source IP
resolvectl query api.example.com # DNS through systemd-resolved
ss -s # socket totals by state
mtr -rwzc 50 api.example.com # loss and latency per hop
curl -so /dev/null -w "dns %{time_namelookup} connect %{time_connect} tls %{time_appconnect} ttfb %{time_starttransfer} total %{time_total}\n" https://api.example.com$ ss -s Total: 1893 TCP: 1712 (estab 412, closed 1201, orphaned 3, timewait 1188) $ curl -so /dev/null -w "..." https://api.example.com dns 0.004 connect 0.031 tls 0.089 ttfb 2.912 total 2.915
The curl timing line is the fastest way to split blame. Here DNS, TCP and TLS take 89 ms together, and the remaining 2.8 s is the server thinking. That is an application or database problem, not the network.
| Symptom | Likely cause | Check |
|---|---|---|
Slow time_namelookup | DNS server slow or unreachable. | resolvectl status, try another resolver with dig @1.1.1.1 |
Slow time_connect | Distance, packet loss, or SYN queue full on the server. | mtr, nstat -az | grep -i listen |
Loss only at the last hop in mtr | Real loss at the destination. | Loss in the middle that does not carry on is usually ICMP rate limiting. |
Thousands in timewait | Many short outbound connections. | Use keepalive and connection pools. |
| Errors and drops on the interface | Bad link, full ring buffer. | ip -s link, ethtool -S eth0 |
Ports and sockets
Who is listening, on which address, and who already holds the port you need?
sudo ss -tulpn # listening TCP and UDP, with process
sudo ss -tlnp "sport = :3000" # one port
sudo lsof -iTCP:3000 -sTCP:LISTEN # same, another way
sudo fuser -v 3000/tcp # which PID holds it
ss -tn state established "dport = :5432" | wc -l # open connections to PostgreSQL
nc -zv 127.0.0.1 6379 # can I connect at all?$ sudo ss -tulpn Netid State Local Address:Port Process tcp LISTEN 0.0.0.0:22 users:(("sshd",pid=702,fd=3)) tcp LISTEN 0.0.0.0:443 users:(("nginx",pid=881,fd=7)) tcp LISTEN 127.0.0.1:3000 users:(("node",pid=2318,fd=20)) tcp LISTEN 127.0.0.1:5432 users:(("postgres",pid=911,fd=6)) $ nc -zv 127.0.0.1 6379 Connection to 127.0.0.1 6379 port [tcp/redis] succeeded!
| Local address | Reachable from |
|---|---|
127.0.0.1:PORT, [::1]:PORT | This machine only. Right for app servers behind nginx and for databases. |
0.0.0.0:PORT, *:PORT | Every IPv4 interface. Only the firewall stands in front of it. |
[::]:PORT | Every IPv6 interface, and often IPv4 too. |
10.0.0.5:PORT | One specific interface, usually the private network. |
Common port errors
| Error | Meaning | Fix |
|---|---|---|
EADDRINUSE, address already in use | Another process holds the port, often an old copy of your app. | sudo ss -tlnp "sport = :3000", stop that process or its unit. |
EACCES on port 80 or 443 | Ports below 1024 need privileges. | Put nginx in front, or AmbientCapabilities=CAP_NET_BIND_SERVICE in the unit. |
ECONNREFUSED locally | Nothing listens on that address and port. | Check the bind address: 127.0.0.1 vs your private IP. |
EADDRNOTAVAIL on outbound | Out of ephemeral ports. | sysctl net.ipv4.ip_local_port_range, reuse connections. |
too many open files | Hit the file descriptor limit; every socket is a file. | cat /proc/PID/limits, raise LimitNOFILE. |
Firewall
The service is listening but clients still cannot reach it. Find which wall drops the packet: local firewall, cloud security group, or the path in between.
Swipe sideways to see the whole diagram
- 1Refused: the packet arrived and the host answered with a reset. Nothing listens there, or a REJECT rule: check module 08.
- 2Timed out: the packet vanished. A DROP rule, a cloud security group, or no route back.
- 3No route to host: an ICMP reject came back, or the host is down on the local network.
Refused, timed out, or no route?
The error a client sees tells you where to look before you open any firewall tool.
| Client sees | What happened | Look at |
|---|---|---|
| Connection refused | The packet arrived and the host answered with a reset: nothing listens, or a REJECT rule. | Module 08, then ufw reject rules. |
| Connection timed out | The packet vanished: a DROP rule, a cloud security group, or no route back. | Firewall rules, cloud console, then tcpdump. |
| No route to host | An ICMP reject came back, or the host is down on the local network. | REJECT rules with icmp-host-prohibited, ARP, ip route. |
| Works locally, not remotely | Bound to 127.0.0.1, or blocked on the way in. | ss -tulpn, then firewall. |
Read the local rules
sudo ufw status numbered # Ubuntu friendly front end
sudo nft list ruleset # the real kernel rules (nftables)
sudo iptables -L -n -v --line-numbers # legacy view, with packet counters
sudo firewall-cmd --list-all # RHEL, Rocky, Fedora (firewalld)
journalctl -k --grep "UFW BLOCK" --since "10 min ago"$ sudo ufw status numbered Status: active To Action From -- ------ ---- [ 1] 22/tcp LIMIT IN Anywhere [ 2] 80,443/tcp ALLOW IN Anywhere $ journalctl -k --grep "UFW BLOCK" -n 1 kernel: [UFW BLOCK] IN=eth0 SRC=203.0.113.9 DST=10.0.0.5 PROTO=TCP SPT=51544 DPT=8080 SYN
The UFW BLOCK line names the source, destination port and protocol of the dropped packet. Here someone tried port 8080, which no rule allows. If the counters on an iptables -v rule climb while you retry, that rule is the one matching.
Prove where the packet stops
tcpdump shows packets after they reach the interface but before most firewall rules apply. No SYN on the wire means the problem is upstream: cloud security group, load balancer, routing. A SYN with no SYN-ACK back means the local host dropped it.
sudo tcpdump -ni any 'tcp port 443 and tcp[tcpflags] & tcp-syn != 0' -c 1014:31:02.114 eth0 In IP 203.0.113.9.51544 > 10.0.0.5.443: Flags [S], seq 1103..., length 0 14:31:02.114 eth0 Out IP 10.0.0.5.443 > 203.0.113.9.51544: Flags [S.], seq 2291..., length 0 # SYN in, SYN-ACK out: the firewall is fine, look at the app
Quick fixes
| Need | ufw | firewalld |
|---|---|---|
| Open a port | sudo ufw allow 8080/tcp | sudo firewall-cmd --add-port=8080/tcp --permanent && sudo firewall-cmd --reload |
| Open to one IP | sudo ufw allow from 10.0.0.12 to any port 5432 proto tcp | --add-rich-rule='rule family=ipv4 source address=10.0.0.12 port port=5432 protocol=tcp accept' |
| Remove a rule | sudo ufw delete 3 | --remove-port=8080/tcp --permanent |
| See blocked packets | sudo ufw logging medium | --set-log-denied=all |
Docker publishes ports by writing its own iptables rules, which bypass ufw. A container started with -p 8080:8080 is public even when ufw denies 8080. Bind it to localhost with -p 127.0.0.1:8080:8080.
Cause found?
A cause is only the cause when it explains the symptom you wrote down in module 01: the timing, the scope and every effect.
| Test | Pass looks like | Fail looks like |
|---|---|---|
| Timing | Memory hit the limit at 14:05; the errors started at 14:05. | The disk was full yesterday too, and nothing broke then. |
| Scope | Only the API on web-1 failed, and only web-1 is out of memory. | Every host is slow but only one has high CPU. |
| Mechanism | You can explain each step from cause to symptom. | "It is probably related." |
| Reproduce | Triggering the cause on staging produces the same symptom. | It cannot be triggered on purpose. |
| Prediction | If you are right, the fix will move a specific number. | There is no number you expect to change. |
If any row fails, treat what you found as a contributing factor, not the cause, and go to module 12. Saturated resources are often an effect: a slow database query makes the app hold connections, which grows memory, which triggers swapping.
Fix and verify
Mitigate first, then fix. Prove it with the same command that showed the problem, and write it down.
| Cause | Mitigate now | Fix for good |
|---|---|---|
| OOM kills | sudo systemctl restart letterpad | Find the leak, set --max-old-space-size, right size MemoryMax. |
| Disk full | sudo journalctl --vacuum-size=500M, restart processes from lsof +L1 | Log retention, alerts at 80%. |
| CPU hot loop | Restart, scale out, rate limit the endpoint. | Profile and fix the code path. |
| Port in use | Stop the stray process. | One unit owns the port; no manual node runs. |
| Firewall drop | sudo ufw allow the exact port and source. | Rule kept in config management, reviewed. |
| Too many open files | Restart. | LimitNOFILE=65535 and fix the socket leak. |
Verify with the same command
systemctl show letterpad -p NRestarts -p ActiveState
journalctl -p err --since "10 min ago" --no-pager | wc -l
curl -so /dev/null -w "%{http_code} %{time_total}\n" https://letterpad.io/health$ systemctl show letterpad -p NRestarts -p ActiveState NRestarts=0 ActiveState=active $ journalctl -p err --since "10 min ago" --no-pager | wc -l 0 $ curl -so /dev/null -w "%{http_code} %{time_total}\n" https://letterpad.io/health 200 0.041
Write it down
| Section | One line each |
|---|---|
| Symptom | The sentence from module 01, with times. |
| Impact | Who was affected, for how long. |
| Cause | The mechanism, from trigger to symptom. |
| Evidence | The commands and outputs that proved it. |
| Fix | What changed, and the before and after numbers. |
| Follow up | The alert or guard that catches it earlier next time. |
Trace deeper
When the usual tools show nothing, watch the process itself: its system calls, its open files, and where it spends CPU.
/proc: everything about one process
PID=2318
cat /proc/$PID/status | head -n 12 # state, threads, memory
cat /proc/$PID/limits # open files, processes, memory limits
ls /proc/$PID/fd | wc -l # how many file descriptors are open
sudo cat /proc/$PID/stack # kernel stack, for a process stuck in D
sudo cat /proc/$PID/environ | tr '\0' '\n' # environment it really gotstrace: what is it waiting for?
sudo apt install -y strace
sudo strace -f -tt -T -p 2318 -e trace=network,read,write -o /tmp/trace.txt
sudo strace -c -f -p 2318 # Ctrl+C after 10s: syscall summary$ sudo strace -c -f -p 2318 % time seconds usecs/call calls errors syscall ------ ----------- ----------- --------- --------- ---------------- 71.40 2.118430 4236 500 500 connect 18.02 0.534660 37 14400 epoll_wait # 500 failed connects: the app cannot reach a dependency
| Flag | Meaning |
|---|---|
-p PID | Attach to a running process. |
-f | Follow threads and child processes. Needed for Node, Java and anything threaded. |
-tt -T | Wall clock timestamps, and time spent in each call. |
-e trace=network | Only show a class of calls: file, network, process, memory. |
-c | A summary table instead of every call. |
strace slows the traced process a lot. Attach for seconds, not minutes, in production.
lsof: what does it have open?
sudo lsof -p 2318 | awk '{print $5}' | sort | uniq -c | sort -rn # count by type
sudo lsof -p 2318 -i # only its sockets, with peersperf: where does the CPU go?
sudo apt install -y linux-tools-common linux-tools-$(uname -r)
sudo perf top -p 2318 # live hottest functions
sudo perf record -F 99 -g -p 2318 -- sleep 30
sudo perf report --stdio | head -n 40For Node: start it with `--perf-basic-prof` so perf can name JavaScript functions instead of showing raw addresses. `node --cpu-prof` writes a profile Chrome DevTools can open, without perf at all.
Which one do I need?
The flowchart as one table. Find the symptom, run the first command.
| Symptom | Area | First command |
|---|---|---|
| Something broke, no idea what | Logs | journalctl -p err -b |
| Service keeps restarting | Service | systemctl status UNIT |
| High load average | CPU | mpstat 1 then pidstat -u 1 |
| Process was killed | Memory | journalctl -k --grep "Out of memory" |
| Box is slow and swapping | Memory | vmstat 1 |
| No space left on device | Disk | df -h; df -i; sudo lsof +L1 |
| Queries and writes are slow | Disk | iostat -xz 1 |
| Requests are slow end to end | Network | curl -w timing line |
| Packet loss or flaky links | Network | mtr -rwzc 50 HOST |
| Address already in use | Ports | sudo ss -tlnp "sport = :PORT" |
| Connection refused | Ports | sudo ss -tulpn |
| Connection timed out | Firewall | sudo ufw status numbered |
| Is the packet even arriving? | Firewall | sudo tcpdump -ni any port PORT |
| Too many open files | Process | cat /proc/PID/limits |
| Stuck, no obvious cause | Trace | sudo strace -c -f -p PID |
| CPU busy in my own code | Trace | sudo perf top -p PID |
Install the toolkit once on every server so it is there when you need it: sudo apt install -y sysstat htop iotop mtr-tiny strace tcpdump lsof.