Linux triage 0/12

Troubleshooting flowchartProcesses, resources, ports, firewallFrom symptom to root cause

Linux
Triage

A server misbehaves and every tool is one command away. This is the order to use them in: pin down the symptom, read the logs, check the service, then test CPU, memory, disk, network, ports and the firewall one at a time until one of them explains what you saw. If none does, trace the process itself.

Linuxjournalctlsystemctlssufwstrace

The flowchartTap a box to jump to its module

01

Define the symptom

Write one sentence before you run a single command: what is wrong, since when, and for whom.

WhatWhenHow widespread

Most wasted hours start with a vague symptom. "The server is slow" sends you everywhere. "Checkout API p95 went from 200 ms to 4 s at 14:05 for all users, other APIs are fine" tells you where to look and what a fix must change. Get the clock, the scope and the last change first.

QuestionWhy it narrows the searchHow to answer it
What exactly fails?Errors, slowness and timeouts lead to different modules.Reproduce it: curl -sv -o /dev/null -w "%{http_code} %{time_total}\n" URL
Since when?Lets you line up logs, deploys and cron jobs.journalctl --since "14:00" --until "14:15"
How widespread?One user, one host, or everyone points at client, node or shared dependency.Check a second host and a second endpoint.
What changed?Most incidents follow a deploy, config edit or package upgrade.grep -E "install|upgrade" /var/log/dpkg.log | tail
Is it still happening?A past spike needs history; a live one needs live tools.uptime, journalctl -f

Sixty seconds of context

Run these first on any box you just logged into. They tell you how long it has been up, how loaded it is, who else is on it, and whether the kernel has complained.

first-minute.shBASH
uptime                     # uptime and load averages
who                        # who else is logged in right now
last -n 5 reboot           # recent reboots
dmesg -T --level=err,warn | tail -n 20
systemctl list-units --failed
TerminalOutput
$ uptime
 14:12:03 up 41 days,  3:10,  2 users,  load average: 7.92, 6.10, 2.41
$ systemctl list-units --failed
  UNIT              LOAD   ACTIVE SUB    DESCRIPTION
● letterpad.service loaded failed failed letterpad.io NestJS API
1 loaded units listed.

Read load averages left to right: the last 1, 5 and 15 minutes. 7.92, 6.10, 2.41 on a 4 core box means it got busy recently and is still climbing. Compare with nproc.

02

Check logs

Errors from this boot first, then narrow by unit, time and text. The kernel log tells you about OOM kills and disk errors.

journalctlKernel/var/log
logs.shBASH
journalctl -p err -b                         # errors and worse since boot
journalctl -u nginx -u letterpad --since "15 min ago"
journalctl -k -b                             # kernel only (same as dmesg)
journalctl -b -1 -e                          # end of the previous boot
journalctl -u letterpad --grep "ECONN|timeout" -n 50
TerminalOutput
$ journalctl -p err -b --no-pager | tail -n 4
Oct 10 14:05:11 web-1 kernel: Out of memory: Killed process 2241 (node) total-vm:2843120kB, anon-rss:1530212kB
Oct 10 14:05:11 web-1 systemd[1]: letterpad.service: A process of this unit has been killed by the OOM killer.
Oct 10 14:05:11 web-1 systemd[1]: letterpad.service: Failed with result 'oom-kill'.
Oct 10 14:05:16 web-1 systemd[1]: letterpad.service: Scheduled restart job, restart counter is at 3.
FlagMeaning
-p errPriority error and worse. Levels: emerg alert crit err warning notice info debug.
-b, -b -1This boot, the previous boot. --list-boots shows them all.
-u UNITOne service. Repeat it to interleave several on one timeline.
-kKernel messages: OOM killer, disk I/O errors, segfaults, link up and down.
--since / --untilA window: "2026-10-10 14:00", "1 hour ago", today.
-f, -e, -n NFollow live, jump to the end, last N lines.
--grepRegex on the message text.
-o json-prettyEvery field, including _PID, _EXE and _SYSTEMD_UNIT.

Files that are not in the journal

FileWhat lands there
/var/log/nginx/*.logAccess and error logs. 502 and 504 explanations live in the error log.
/var/log/auth.logSSH logins, sudo use, failed passwords.
/var/log/syslogUbuntu's plain text copy of most of the journal.
/var/log/postgresql/Slow queries, lock waits, connection limit errors.
/var/log/apt/history.logWhat was installed or upgraded, and when.
03

Check service state

Is the process running, restarting, stuck or a zombie? systemd tells you about the unit; ps tells you about every process.

systemctlpsSignals
state.shBASH
systemctl status letterpad
systemctl show letterpad -p ActiveState -p SubState -p NRestarts -p Result
ps -eo pid,ppid,user,stat,%cpu,%mem,etime,cmd --sort=-%cpu | head
pstree -p $(systemctl show -p MainPID --value letterpad)
pgrep -a node
TerminalOutput
$ ps -eo pid,ppid,user,stat,%cpu,%mem,etime,cmd --sort=-%cpu | head -n 4
    PID    PPID USER      STAT %CPU %MEM     ELAPSED CMD
   2318       1 letterpad Rsl  97.4 18.2       02:11 /usr/bin/node dist/main.js
    911       1 postgres  Ss    3.1  2.4 41-03:09:55 /usr/lib/postgresql/16/bin/postgres
   1312       1 rabbitmq  Ssl   1.0  3.8 41-03:09:51 /usr/lib/erlang/erts/bin/beam.smp

Read the STAT column

StateMeaningWhat to do
RRunning or ready to run.Normal. Many R processes plus high load means CPU pressure.
SSleeping, waiting for an event.Normal for idle servers.
DUninterruptible sleep, usually waiting on disk or NFS.Cannot be killed. Look at module 06. Many D processes push load up without CPU use.
ZZombie: exited, parent has not collected it.Harmless alone. Thousands mean the parent is buggy; restart the parent.
TStopped by a signal or a debugger.kill -CONT PID resumes it.
s l + < NSession leader, multi-threaded, foreground, high and low priority.Extra flags after the state letter.

Signals you will actually send

CommandSignalEffect
kill PIDTERM (15)Ask to exit cleanly. Always try first.
kill -HUP PIDHUP (1)Many daemons reload config on it (nginx, sshd).
kill -INT PIDINT (2)Same as Ctrl+C.
kill -9 PIDKILL (9)Immediate, no cleanup. Last resort, and useless on D state.
kill -STOP / -CONTSTOP, CONTPause and resume.
pkill -f "dist/main.js"TERMMatch by full command line. Check with pgrep -af first.

For a service under systemd, prefer systemctl restart over kill: systemd knows the whole process tree and the restart policy.

04

Resources: CPU

Check resources one at a time and stop at the first one that is saturated. Start with CPU: who is using it, and is it really user code?

toppidstatSteal and iowait
Where the CPU time goes

Swipe sideways to see the whole diagram

mpstat 1 1: all CPUs, one second%usr 88.2: your code%sys 6.0idle 4.8pidstat names the processnode, PID 231895% CPU, running on core 2
  1. 1%usr near 90 means your own code is busy. Find the hot path with a profiler, module 12.
  2. 2A high %iowait would mean the CPU sits idle waiting on the disk: go to module 06 instead.
  3. 3pidstat names the process: here node, PID 2318, at 95% of one core.

For each resource ask three things, the USE method: Utilisation (how busy), Saturation (is work queuing) and Errors. Install sysstat once to get pidstat, mpstat and iostat.

cpu.shBASH
sudo apt install -y sysstat htop
nproc                       # number of CPUs, to compare with load
top -o %CPU                 # press 1 for per-CPU, P sort by CPU, H threads
mpstat -P ALL 1 3           # per-CPU breakdown, 3 samples
pidstat -u 1 5              # per-process CPU every second
pidstat -t -p 2318 1        # per-thread for one PID
TerminalOutput
$ mpstat 1 1
CPU    %usr   %nice    %sys %iowait    %irq   %soft  %steal  %idle
all   88.21    0.00    6.02    0.25    0.00    0.75    0.00    4.77
$ pidstat -u 1 1
   UID       PID    %usr %system  %CPU   CPU  Command
   998      2318   92.00    3.00  95.00     2  node
ColumnHigh value means
%usrYour code is busy: a hot loop, JSON work, regex, crypto. Profile it (module 12).
%sysThe kernel is busy for you: many syscalls, context switches, small writes.
%iowaitCPU idle while waiting for disk. It is a disk problem, go to module 06.
%stealThe hypervisor gave your CPU to another VM. Nothing to fix inside; resize or move.
%softPacket processing. Very high under heavy network load.
load average > nprocWork is queuing. Includes D state processes, so check iowait before blaming CPU.

Node.js runs your JavaScript on one thread, so one node at 100% on a 4 core box is saturated even though top shows 25% overall. Watch per-process numbers, not just the total.

05

Resources: memory

Look at available, not free. Then check swap activity and whether the OOM killer has been busy.

freevmstatOOM killer
Available, not free

Swipe sideways to see the whole diagram

RAM 3.8Giused 3.5Giavailable 161Miused 3.5Gibuff/cache 240Mifree 112MiSwap 2.0Gi1.9Gi usedpages out, so 2210 KB/spages in, si 1840 KB/sLittle available and busy swap: real memory pressure.free on its own misleads: Linux keeps spare RAM as cache on purpose.vmstat agrees: 4 processes blocked on I/O in the b column.
  1. 1Read available, not free. Linux fills spare RAM with cache and hands it back when asked.
  2. 2si and so above zero for seconds means the box is swapping right now, and everything slows down.
  3. 3For a Node service, cap the heap below the unit's MemoryMax with --max-old-space-size, so you get a heap error you can read instead of the OOM killer.
memory.shBASH
free -h
vmstat 1 5                  # si/so = swap in/out per second
ps -eo pid,user,rss,vsz,comm --sort=-rss | head
journalctl -k --grep "Out of memory|oom-kill" --since today
cat /proc/2318/status | grep -E "VmRSS|VmSwap|Threads"
systemctl show letterpad -p MemoryCurrent -p MemoryMax
TerminalOutput
$ free -h
               total        used        free      shared  buff/cache   available
Mem:           3.8Gi       3.5Gi       112Mi        24Mi       240Mi       161Mi
Swap:          2.0Gi       1.9Gi       100Mi
$ vmstat 1 2
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
 r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st
 3  4 1992344 114820  8204 237600 1840 2210  9120  2380 3110 5902 21 14  9 56  0
SignalWhat it tells you
available near zeroReal memory pressure. free alone is misleading: Linux uses spare RAM as cache on purpose.
si / so above zero for secondsThe box is swapping right now. Everything gets slow and wa rises.
b column above zeroProcesses blocked on I/O, often caused by swapping.
Killed process in the kernel logThe OOM killer picked a victim. The log line names it and its RSS.
RSS of one process keeps growingA leak. Compare ps a few minutes apart, or watch MemoryCurrent.

For a Node service, cap the heap below the unit's MemoryMax with NODE_OPTIONS=--max-old-space-size=512, so Node throws a heap error you can read instead of the kernel killing it silently.

06

Resources: disk

Two different problems hide here: a disk that is full (space or inodes) and a disk that is slow (latency and queueing).

dfiostatDeleted but open
Full, or slow?

Swipe sideways to see the whole diagram

Is it full?deleted log 8.5G48G of 48G100% used/dev/vda1, ext4lsof +L1: node 2318 holds itIs it slow?aqu-sz 48: requests waitingvda%util 99.6w_await 41.9 ms per writer_await 18.4 ms per readAn SSD should be single digit.
  1. 1Full: df says 100%. A log removed with rm while node still has it open keeps its 8.5G until that process lets go.
  2. 2Free it by restarting the process, or truncate it in place with : > /proc/2318/fd/21.
  3. 3Slow: aqu-sz above 1 for long means requests are waiting. On NVMe and RAID, trust await over %util.

Is it full?

disk-space.shBASH
df -hT -x tmpfs -x devtmpfs       # space per filesystem
df -i                             # inodes: full even with free space
sudo du -xh --max-depth=1 /var | sort -h | tail
sudo lsof +L1                     # deleted files still held open
TerminalOutput
$ df -hT -x tmpfs -x devtmpfs
Filesystem     Type  Size  Used Avail Use% Mounted on
/dev/vda1      ext4   48G   48G     0 100% /
$ sudo lsof +L1
COMMAND  PID     USER  FD  TYPE DEVICE   SIZE/OFF NLINK  NODE NAME
node    2318 letterpad 21w REG  252,1 9126805504     0 51234 /var/lib/letterpad/debug.log (deleted)

lsof +L1 catches the classic trap: someone deleted a huge log with rm, but the process still has it open, so the space is not freed. Restart that process, or truncate it in place with : > /proc/2318/fd/21.

Is it slow?

disk-io.shBASH
iostat -xz 1 3                    # extended stats, skip idle devices
pidstat -d 1                      # read/write per process
sudo iotop -oPa                   # live, only processes doing I/O
TerminalOutput
$ iostat -xz 1 1
Device   r/s    w/s   rkB/s   wkB/s  r_await w_await aqu-sz  %util
vda    412.0  880.0 26368.0 70400.0   18.40   41.90  48.21  99.60
ColumnRead it as
r_await / w_awaitAverage ms per request, queue time included. SSDs should be single digit.
aqu-szAverage queue length. Above 1 for long means requests are waiting.
%utilTime the device was busy. Near 100 on a single disk means saturated; on NVMe and RAID it can mislead, trust await.
07

Resources: network

Work outward: is the interface up, is there a route, does DNS resolve, does the path drop packets, and where does a request spend its time?

ipss -smtr
Where 2.9 seconds went

Swipe sideways to see the whole diagram

DNS4 msConnect27 msTLS58 msServer2,823 ms: the server thinkingDownload3 ms01 s2 s3 sNetwork: 89 ms. Server: 2.8 s. Fix the app, not the network.
  1. 1The curl -w times count from the start, so subtract each from the next to get one phase.
  2. 2A slow time_namelookup points at DNS; a slow time_connect at distance, packet loss or a full SYN queue.
  3. 3Here DNS, TCP and TLS take 89 ms together. The other 2.8 s is the application or the database.
network.shBASH
ip -br addr                        # interfaces and addresses
ip route get 1.1.1.1               # which route and source IP
resolvectl query api.example.com   # DNS through systemd-resolved
ss -s                              # socket totals by state
mtr -rwzc 50 api.example.com       # loss and latency per hop
curl -so /dev/null -w "dns %{time_namelookup} connect %{time_connect} tls %{time_appconnect} ttfb %{time_starttransfer} total %{time_total}\n" https://api.example.com
TerminalOutput
$ ss -s
Total: 1893
TCP:   1712 (estab 412, closed 1201, orphaned 3, timewait 1188)
$ curl -so /dev/null -w "..." https://api.example.com
dns 0.004 connect 0.031 tls 0.089 ttfb 2.912 total 2.915

The curl timing line is the fastest way to split blame. Here DNS, TCP and TLS take 89 ms together, and the remaining 2.8 s is the server thinking. That is an application or database problem, not the network.

SymptomLikely causeCheck
Slow time_namelookupDNS server slow or unreachable.resolvectl status, try another resolver with dig @1.1.1.1
Slow time_connectDistance, packet loss, or SYN queue full on the server.mtr, nstat -az | grep -i listen
Loss only at the last hop in mtrReal loss at the destination.Loss in the middle that does not carry on is usually ICMP rate limiting.
Thousands in timewaitMany short outbound connections.Use keepalive and connection pools.
Errors and drops on the interfaceBad link, full ring buffer.ip -s link, ethtool -S eth0
08

Ports and sockets

Who is listening, on which address, and who already holds the port you need?

ss -tulpnlsof -iBind address
ports.shBASH
sudo ss -tulpn                      # listening TCP and UDP, with process
sudo ss -tlnp "sport = :3000"       # one port
sudo lsof -iTCP:3000 -sTCP:LISTEN   # same, another way
sudo fuser -v 3000/tcp              # which PID holds it
ss -tn state established "dport = :5432" | wc -l   # open connections to PostgreSQL
nc -zv 127.0.0.1 6379               # can I connect at all?
TerminalOutput
$ sudo ss -tulpn
Netid State  Local Address:Port  Process
tcp   LISTEN 0.0.0.0:22          users:(("sshd",pid=702,fd=3))
tcp   LISTEN 0.0.0.0:443         users:(("nginx",pid=881,fd=7))
tcp   LISTEN 127.0.0.1:3000      users:(("node",pid=2318,fd=20))
tcp   LISTEN 127.0.0.1:5432      users:(("postgres",pid=911,fd=6))
$ nc -zv 127.0.0.1 6379
Connection to 127.0.0.1 6379 port [tcp/redis] succeeded!
Local addressReachable from
127.0.0.1:PORT, [::1]:PORTThis machine only. Right for app servers behind nginx and for databases.
0.0.0.0:PORT, *:PORTEvery IPv4 interface. Only the firewall stands in front of it.
[::]:PORTEvery IPv6 interface, and often IPv4 too.
10.0.0.5:PORTOne specific interface, usually the private network.

Common port errors

ErrorMeaningFix
EADDRINUSE, address already in useAnother process holds the port, often an old copy of your app.sudo ss -tlnp "sport = :3000", stop that process or its unit.
EACCES on port 80 or 443Ports below 1024 need privileges.Put nginx in front, or AmbientCapabilities=CAP_NET_BIND_SERVICE in the unit.
ECONNREFUSED locallyNothing listens on that address and port.Check the bind address: 127.0.0.1 vs your private IP.
EADDRNOTAVAIL on outboundOut of ephemeral ports.sysctl net.ipv4.ip_local_port_range, reuse connections.
too many open filesHit the file descriptor limit; every socket is a file.cat /proc/PID/limits, raise LimitNOFILE.
09

Firewall

The service is listening but clients still cannot reach it. Find which wall drops the packet: local firewall, cloud security group, or the path in between.

ufwnftablestcpdump
Refused, timed out, or no route

Swipe sideways to see the whole diagram

ClientHostfirewallRefusednothing listens: RSTTimed outDROP rule: silencewaits, then gives upNo route to hostREJECT: ICMP back
  1. 1Refused: the packet arrived and the host answered with a reset. Nothing listens there, or a REJECT rule: check module 08.
  2. 2Timed out: the packet vanished. A DROP rule, a cloud security group, or no route back.
  3. 3No route to host: an ICMP reject came back, or the host is down on the local network.

Refused, timed out, or no route?

The error a client sees tells you where to look before you open any firewall tool.

Client seesWhat happenedLook at
Connection refusedThe packet arrived and the host answered with a reset: nothing listens, or a REJECT rule.Module 08, then ufw reject rules.
Connection timed outThe packet vanished: a DROP rule, a cloud security group, or no route back.Firewall rules, cloud console, then tcpdump.
No route to hostAn ICMP reject came back, or the host is down on the local network.REJECT rules with icmp-host-prohibited, ARP, ip route.
Works locally, not remotelyBound to 127.0.0.1, or blocked on the way in.ss -tulpn, then firewall.

Read the local rules

firewall.shBASH
sudo ufw status numbered            # Ubuntu friendly front end
sudo nft list ruleset               # the real kernel rules (nftables)
sudo iptables -L -n -v --line-numbers   # legacy view, with packet counters
sudo firewall-cmd --list-all        # RHEL, Rocky, Fedora (firewalld)
journalctl -k --grep "UFW BLOCK" --since "10 min ago"
TerminalOutput
$ sudo ufw status numbered
Status: active
     To                         Action      From
     --                         ------      ----
[ 1] 22/tcp                     LIMIT IN    Anywhere
[ 2] 80,443/tcp                 ALLOW IN    Anywhere
$ journalctl -k --grep "UFW BLOCK" -n 1
kernel: [UFW BLOCK] IN=eth0 SRC=203.0.113.9 DST=10.0.0.5 PROTO=TCP SPT=51544 DPT=8080 SYN

The UFW BLOCK line names the source, destination port and protocol of the dropped packet. Here someone tried port 8080, which no rule allows. If the counters on an iptables -v rule climb while you retry, that rule is the one matching.

Prove where the packet stops

tcpdump shows packets after they reach the interface but before most firewall rules apply. No SYN on the wire means the problem is upstream: cloud security group, load balancer, routing. A SYN with no SYN-ACK back means the local host dropped it.

tcpdump.shBASH
sudo tcpdump -ni any 'tcp port 443 and tcp[tcpflags] & tcp-syn != 0' -c 10
TerminalOutput
14:31:02.114 eth0 In  IP 203.0.113.9.51544 > 10.0.0.5.443: Flags [S], seq 1103..., length 0
14:31:02.114 eth0 Out IP 10.0.0.5.443 > 203.0.113.9.51544: Flags [S.], seq 2291..., length 0
# SYN in, SYN-ACK out: the firewall is fine, look at the app

Quick fixes

Needufwfirewalld
Open a portsudo ufw allow 8080/tcpsudo firewall-cmd --add-port=8080/tcp --permanent && sudo firewall-cmd --reload
Open to one IPsudo ufw allow from 10.0.0.12 to any port 5432 proto tcp--add-rich-rule='rule family=ipv4 source address=10.0.0.12 port port=5432 protocol=tcp accept'
Remove a rulesudo ufw delete 3--remove-port=8080/tcp --permanent
See blocked packetssudo ufw logging medium--set-log-denied=all

Docker publishes ports by writing its own iptables rules, which bypass ufw. A container started with -p 8080:8080 is public even when ufw denies 8080. Bind it to localhost with -p 127.0.0.1:8080:8080.

10

Cause found?

A cause is only the cause when it explains the symptom you wrote down in module 01: the timing, the scope and every effect.

TimingScopeEvidence
TestPass looks likeFail looks like
TimingMemory hit the limit at 14:05; the errors started at 14:05.The disk was full yesterday too, and nothing broke then.
ScopeOnly the API on web-1 failed, and only web-1 is out of memory.Every host is slow but only one has high CPU.
MechanismYou can explain each step from cause to symptom."It is probably related."
ReproduceTriggering the cause on staging produces the same symptom.It cannot be triggered on purpose.
PredictionIf you are right, the fix will move a specific number.There is no number you expect to change.

If any row fails, treat what you found as a contributing factor, not the cause, and go to module 12. Saturated resources are often an effect: a slow database query makes the app hold connections, which grows memory, which triggers swapping.

11

Fix and verify

Mitigate first, then fix. Prove it with the same command that showed the problem, and write it down.

MitigateVerifyDocument
CauseMitigate nowFix for good
OOM killssudo systemctl restart letterpadFind the leak, set --max-old-space-size, right size MemoryMax.
Disk fullsudo journalctl --vacuum-size=500M, restart processes from lsof +L1Log retention, alerts at 80%.
CPU hot loopRestart, scale out, rate limit the endpoint.Profile and fix the code path.
Port in useStop the stray process.One unit owns the port; no manual node runs.
Firewall dropsudo ufw allow the exact port and source.Rule kept in config management, reviewed.
Too many open filesRestart.LimitNOFILE=65535 and fix the socket leak.

Verify with the same command

verify.shBASH
systemctl show letterpad -p NRestarts -p ActiveState
journalctl -p err --since "10 min ago" --no-pager | wc -l
curl -so /dev/null -w "%{http_code} %{time_total}\n" https://letterpad.io/health
TerminalOutput
$ systemctl show letterpad -p NRestarts -p ActiveState
NRestarts=0
ActiveState=active
$ journalctl -p err --since "10 min ago" --no-pager | wc -l
0
$ curl -so /dev/null -w "%{http_code} %{time_total}\n" https://letterpad.io/health
200 0.041

Write it down

SectionOne line each
SymptomThe sentence from module 01, with times.
ImpactWho was affected, for how long.
CauseThe mechanism, from trigger to symptom.
EvidenceThe commands and outputs that proved it.
FixWhat changed, and the before and after numbers.
Follow upThe alert or guard that catches it earlier next time.
12

Trace deeper

When the usual tools show nothing, watch the process itself: its system calls, its open files, and where it spends CPU.

stracelsofperf

/proc: everything about one process

proc.shBASH
PID=2318
cat /proc/$PID/status | head -n 12   # state, threads, memory
cat /proc/$PID/limits                # open files, processes, memory limits
ls /proc/$PID/fd | wc -l             # how many file descriptors are open
sudo cat /proc/$PID/stack            # kernel stack, for a process stuck in D
sudo cat /proc/$PID/environ | tr '\0' '\n'   # environment it really got

strace: what is it waiting for?

strace.shBASH
sudo apt install -y strace
sudo strace -f -tt -T -p 2318 -e trace=network,read,write -o /tmp/trace.txt
sudo strace -c -f -p 2318            # Ctrl+C after 10s: syscall summary
TerminalOutput
$ sudo strace -c -f -p 2318
% time     seconds  usecs/call     calls    errors syscall
------ ----------- ----------- --------- --------- ----------------
 71.40    2.118430        4236       500       500 connect
 18.02    0.534660          37     14400           epoll_wait
# 500 failed connects: the app cannot reach a dependency
FlagMeaning
-p PIDAttach to a running process.
-fFollow threads and child processes. Needed for Node, Java and anything threaded.
-tt -TWall clock timestamps, and time spent in each call.
-e trace=networkOnly show a class of calls: file, network, process, memory.
-cA summary table instead of every call.

strace slows the traced process a lot. Attach for seconds, not minutes, in production.

lsof: what does it have open?

lsof.shBASH
sudo lsof -p 2318 | awk '{print $5}' | sort | uniq -c | sort -rn   # count by type
sudo lsof -p 2318 -i                  # only its sockets, with peers

perf: where does the CPU go?

perf.shBASH
sudo apt install -y linux-tools-common linux-tools-$(uname -r)
sudo perf top -p 2318                 # live hottest functions
sudo perf record -F 99 -g -p 2318 -- sleep 30
sudo perf report --stdio | head -n 40

For Node: start it with `--perf-basic-prof` so perf can name JavaScript functions instead of showing raw addresses. `node --cpu-prof` writes a profile Chrome DevTools can open, without perf at all.

Which one do I need?

The flowchart as one table. Find the symptom, run the first command.

Cheat sheet
SymptomAreaFirst command
Something broke, no idea whatLogsjournalctl -p err -b
Service keeps restartingServicesystemctl status UNIT
High load averageCPUmpstat 1 then pidstat -u 1
Process was killedMemoryjournalctl -k --grep "Out of memory"
Box is slow and swappingMemoryvmstat 1
No space left on deviceDiskdf -h; df -i; sudo lsof +L1
Queries and writes are slowDiskiostat -xz 1
Requests are slow end to endNetworkcurl -w timing line
Packet loss or flaky linksNetworkmtr -rwzc 50 HOST
Address already in usePortssudo ss -tlnp "sport = :PORT"
Connection refusedPortssudo ss -tulpn
Connection timed outFirewallsudo ufw status numbered
Is the packet even arriving?Firewallsudo tcpdump -ni any port PORT
Too many open filesProcesscat /proc/PID/limits
Stuck, no obvious causeTracesudo strace -c -f -p PID
CPU busy in my own codeTracesudo perf top -p PID

Install the toolkit once on every server so it is there when you need it: sudo apt install -y sysstat htop iotop mtr-tiny strace tcpdump lsof.