Take a Node API and its frontend from your machine to production: a Linode you can trust, Nginx at the front door, PM2 reloading without dropping a request, releases you can roll back in a minute, and the AWS pieces you actually use, RDS, Lambda, S3 and CloudWatch, with logs and alarms that tell you before users do.
By Shree Kumar Sharma, with Claude Design, for Backend Engineering. Progress is saved in this browser only.
A catalog of every term in this guide: what it means in plain words, what it means technically, and an everyday comparison, each linked to the module that teaches it.
In detail
Filter by kind or type a word to find it. Terms such as reverse proxy, upstream, cluster mode or security group are also linked from the code examples, so you can always jump from a config line to its meaning.
Deployment
Basics
In plain words
Putting your app where users can reach it.
Technically
Building an artifact, shipping it to a runtime, configuring it and switching traffic to it.
Think of it as
Moving a restaurant from your home kitchen into a real shop.
Deploying means taking code that runs on your machine and running it on a computer that is always on, reachable by name, secure, observable and easy to update without breaking users.
In detail
A deployment has five parts: a build (compile TypeScript, bundle the frontend), an artifact (the files you ship), a runtime (Node behind a process manager, or a Lambda), configuration (environment variables and secrets, never in git) and a release process (copy, switch, verify, roll back). Production differs from development in three ways that bite: no hot reload, real secrets, and many users at once. Every module after this one hardens one of those parts.
Build once, ship the artifact, switch trafficClientServiceData
deploy.shShell
npm ci && npm run build # buildrsync -a dist/ server:/srv/app/ # shipssh server "pm2 reload app"# release
#!/usr/bin/env bash# The whole idea of a deployment in twenty linesset -euo pipefailAPP=apiHOST=deploy@203.0.113.10RELEASE=$(date +%Y%m%d%H%M%S)# 1. Build an artifact on a clean machine (CI or your laptop)npm cinpm run build # TypeScript -> dist/, React -> web/dist/# 2. Ship it to a new, versioned folderrsync -az --delete dist/ package.json package-lock.json "$HOST:/srv/$APP/releases/$RELEASE/"# 3. Install production dependencies and switch atomicallyssh"$HOST" bash -s <<EOF cd /srv/$APP/releases/$RELEASE && npm ci --omit=dev ln -sfn /srv/$APP/releases/$RELEASE /srv/$APP/current pm2 reload /srv/$APP/ecosystem.config.js --update-envEOF# 4. Verify, or roll back to the previous releasecurl -fsS https://api.example.com/health || echo "health check failed, roll back"
Why it matters Every later module hardens one of these four steps: build, ship, switch and verify.
A visitor types a domain; DNS returns an IP; the browser connects over TLS to Nginx on port 443; Nginx serves the frontend files itself or forwards API calls to Node on localhost:3000, managed by PM2; Node talks to RDS and logs everything.
In detail
Knowing each hop tells you where to look when something breaks. DNS problems show as 'site not found', TLS problems as certificate warnings, Nginx problems as 502 or 504, app problems as 500s in PM2 logs, and database problems as slow responses or connection errors. In AWS the same path may start with CloudFront or an Application Load Balancer, and an API route may land on Lambda instead of a server.
Every hop is a place to look when it breaksClientExternalEdgeDataService
BrowserClientDNSExternalNginx :443EdgeStatic filesDataNode via PM2ServiceRDS PostgreSQLDataCloudWatchExternallookupHTTPS//apiSQLlogs
# Walk the request path hop by hop when something breaksdig +short app.example.com # 1. DNS: does the name resolve to your IP?nc -zv app.example.com 443# 2. Network: is port 443 open?openssl s_client -connect app.example.com:443 -servername app.example.com </dev/null | head -20# 3. TLS: valid certificate for this name?curl -sI https://app.example.com # 4. Nginx: status and Server headercurl -s http://127.0.0.1:3000/health # 5. On the server: is Node itself healthy?pm2 status # 6. Process: online, restarts, memorysudo tail -n50 /var/log/nginx/error.log # 7. Proxy errors: upstream refused or timed out
Why it matters Each command isolates one hop, so you find the broken layer in minutes instead of guessing.
You can run an app on a server you manage (Linode, EC2), a platform that manages it for you (Elastic Beanstalk, Render), containers (ECS, Kubernetes) or serverless functions (Lambda). Each trades control for convenience and cost shape.
In detail
A VPS is cheap, predictable and teaches you everything, but patching and uptime are on you. PaaS hides servers but limits customisation. Containers give portable, reproducible runtimes at the cost of orchestration. Serverless bills per request and scales to zero, ideal for spiky APIs, cron jobs and event handlers, but has cold starts, time limits and different connection patterns for databases. Many real systems mix them: a VPS or EC2 for the main API behind Nginx, Lambda for background jobs, RDS for data.
Keep code identical across development, staging and production, and change behaviour only through environment variables. Secrets such as database passwords never go in git.
In detail
The Twelve Factor rule: config lives in the environment. Locally use a .env file ignored by git; on a server load variables through PM2's ecosystem file or a systemd EnvironmentFile readable only by the app user; on AWS read them from SSM Parameter Store or Secrets Manager. Validate config at startup and refuse to boot if something is missing. Frontend builds bake public variables in at build time, so never put secrets in them.
config.jsJavaScript
const required = ["DATABASE_URL", "JWT_SECRET"];for (const k of required) if (!process.env[k]) thrownewError(`Missing ${k}`);
Buy a domain, then create DNS records: an A record points a name to your server's IPv4, AAAA to IPv6, CNAME aliases one name to another, and TXT proves ownership for certificates and email.
In detail
Typical setup: example.com and www as A records to your Linode IP, api.example.com to the same server or an AWS load balancer, and the apex on AWS using an alias record to CloudFront. Lower the TTL to 300 seconds a day before a migration so changes spread fast, then raise it again. Route 53, Linode DNS Manager or Cloudflare all work; keep DNS somewhere separate from your server so a dead server never takes your DNS with it.
zone.txtNotes
example.com. 300 A 203.0.113.10www.example.com. 300 CNAME example.com.
; A small production zoneexample.com. 300 A 203.0.113.10 ; Linode VPS running Nginxwww.example.com. 300 CNAME example.com.api.example.com. 300 A 203.0.113.10 ; same box, Nginx routes by server_namestatic.example.com. 300 CNAME d111111abcdef8.cloudfront.net. ; frontend on S3 + CloudFrontexample.com. 300 CAA 0 issue "letsencrypt.org" ; only this CA may issue certsexample.com. 300 TXT "v=spf1 include:amazonses.com ~all"; Check from anywhere; dig +short api.example.com; dig +trace example.com
Why it matters A CAA record is a one line security win: no other certificate authority can issue a certificate for your domain.
02
Phase 02, modules 06 to 11
Your first server
A Linode VPS from zero
Create a server, lock it down, install Node and lay out your app so deploys stay boring.
Pick a region near your users, an Ubuntu LTS image and a shared CPU plan, add your SSH public key, and you have a server with a public IP in about a minute.
In detail
A 2 GB shared Linode comfortably runs Nginx, a Node API in PM2 cluster mode with two workers and a small Redis. Use a dedicated CPU plan when CPU stays above about 50 percent. Enable Linode Backups, attach a Cloud Firewall, and add a private IP if you will connect to another Linode. The same steps map to an EC2 instance on AWS: AMI instead of image, security group instead of Cloud Firewall, Elastic IP for a fixed address.
#!/usr/bin/env bash# Create a server from the terminal with the Linode CLIlinode-cli linodes create \ --label api-prod-1 \ --region ap-south \ --type g6-standard-1 \ --image linode/ubuntu24.04 \ --root_pass"$(openssl rand -base64 32)" \ --authorized_keys"$(cat ~/.ssh/id_ed25519.pub)" \ --backups_enabled true \ --private_ip true \ --tags production# Find its public IPlinode-cli linodes list --label api-prod-1 --format ipv4 --text --no-headers# First login with your key, never a passwordssh root@203.0.113.10
Why it matters Passing your SSH key at creation means password login is never needed, even for the first connection.
Create a non root user with sudo, log in with SSH keys only, disable root login and password authentication, and add fail2ban to block brute force attempts.
In detail
Bots start guessing passwords within minutes of a server going online. Keys make guessing pointless. A separate deploy user owns the app directory and runs PM2, so a compromised app cannot touch system files. Keep a second key or the provider's console (Linode Lish, EC2 Instance Connect) as a way back in before you restart sshd with new settings.
Firewalls: ufw, Cloud Firewall and security groups
Allow only SSH, HTTP and HTTPS from the internet. Node, Redis and PostgreSQL ports stay closed to the world and listen on localhost or a private network.
In detail
Use two layers. A provider firewall (Linode Cloud Firewall, AWS security group) drops traffic before it reaches the server. A host firewall (ufw) protects the box even if the provider rule is wrong. Restrict SSH to your office or VPN IP when you can. On AWS, security groups reference each other: the RDS group allows port 5432 only from the app servers' group, never from 0.0.0.0/0.
Two layers outside, loopback insideClientEdgeServiceData
InternetClientCloud firewallEdgeufw on hostEdgeNginx 80/443ServiceNode 127.0.0.1:3000ServicePostgres privateData22, 80, 443allowedloopback
#!/usr/bin/env bashset -euo pipefailufw default deny incomingufw default allow outgoingufw allow from 198.51.100.0/24 to any port 22 proto tcp # SSH only from office or VPNufw allow "Nginx Full"# 80 and 443ufw --force enableufw status verbose# Prove the app port is NOT reachable from outside# From your laptop: nc -zv 203.0.113.10 3000 -> should time out# And make Node listen on loopback only# app.listen(3000, "127.0.0.1")
Why it matters Binding Node to 127.0.0.1 means even a firewall mistake cannot expose it; only Nginx can reach it.
Install the current Node LTS from NodeSource or with nvm for the deploy user, plus Git and build essentials for native modules. Pin the version so every server and CI run the same one.
In detail
Distribution packages are often years behind. NodeSource gives a system wide Node managed by apt; nvm gives per user versions and is handy when several apps need different ones, but then PM2 must be started from that user's shell so it finds the same node. Write the version in .nvmrc and in package.json engines, and make CI use it too.
#!/usr/bin/env bashset -euo pipefail# System wide Node 22 LTS from NodeSourcecurl -fsSL https://deb.nodesource.com/setup_22.x | sudo -E bash -sudo apt-get install -y nodejs git build-essentialnode -v && npm -v# PM2 for the deploy usersudo npm install -g pm2@latest# Pin the version in the repo so CI and servers agreeecho "22" > .nvmrcnpm pkg set engines.node=">=22 <23"
Why it matters Pinning Node in .nvmrc and engines catches a mismatched runtime before it becomes a production only bug.
Deploy each version into its own timestamped folder under releases/, keep shared files such as .env and uploads in shared/, and point a current symlink at the active release. Switching the symlink is instant and rolling back is just pointing it back.
In detail
Never git pull into the live folder: half updated files serve broken pages mid deploy. With releases, the new version is fully installed and tested before it becomes current. Keep the last five releases for quick rollbacks and delete older ones. PM2 and Nginx both point at /srv/app/current, so neither config changes between releases.
/srv/api/ releases/ 20261004T0930/ older, kept for rollback 20261005T1802/ 20261006T1015/ newest, fully installed before switch shared/ .env secrets, chmod 600, owned by deploy uploads/ user files that must survive deploys logs/ current -> releases/20261006T1015 what PM2 and Nginx point at ecosystem.config.js cwd: /srv/api/current# Switch atomically: rename over the old symlink in one stepln -sfn /srv/api/releases/20261006T1015 /srv/api/current.tmpmv -Tf /srv/api/current.tmp /srv/api/current
Why it matters mv -T replaces the symlink in a single system call, so there is never a moment with no current release.
Back up three things: the server (Linode Backups or EC2 snapshots), the database (RDS automated backups plus manual snapshots before big changes) and user files (S3 with versioning). A backup is only real once you have restored it.
In detail
Follow 3 2 1: three copies, two kinds of storage, one off site. Server images let you rebuild fast but the database is what really matters. RDS keeps point in time recovery for up to 35 days. Write the restore steps down and rehearse them every quarter, timing how long it takes; that time is your real recovery time objective.
Nginx is a fast web server and reverse proxy. It terminates HTTPS, serves static files, forwards API calls to Node and protects it from slow clients and bursts.
In detail
Config lives in /etc/nginx: nginx.conf holds global settings, each site gets a file in sites-available linked into sites-enabled (or a file in conf.d). Blocks nest: http contains server blocks (one per domain), which contain location blocks (one per path). Always run nginx -t before reloading. Reload is graceful: old workers finish their requests while new ones take over.
# /etc/nginx/nginx.conf (trimmed)user www-data;worker_processes auto; # one per CPU coreevents { worker_connections 4096; }http { include /etc/nginx/mime.types; sendfile on; tcp_nopush on; keepalive_timeout65; server_tokens off; # hide the version number client_max_body_size 10m; # upload limit include /etc/nginx/conf.d/*.conf; include /etc/nginx/sites-enabled/*;}# Commands you will use every day# sudo nginx -t test config# sudo systemctl reload nginx graceful reload# sudo ln -s /etc/nginx/sites-available/app /etc/nginx/sites-enabled/
Why it matters worker_processes auto and a test before every reload are the two habits that prevent most Nginx outages.
A location block with proxy_pass forwards API requests to Node on localhost. Pass the original host, client IP and protocol in headers so the app can log and build URLs correctly.
In detail
Set Host, X-Real-IP, X-Forwarded-For and X-Forwarded-Proto, and tell Express to trust the proxy so req.ip and req.protocol are right. Keep connections to Node alive with an upstream keepalive pool. Tune timeouts: proxy_read_timeout of 60 seconds suits normal APIs, longer for slow exports. 502 means Node refused or crashed; 504 means it was too slow.
TLS ends at Nginx; Node speaks plain HTTP on loopbackClientEdgeService
Point Nginx's root at the built frontend folder, serve files directly, and fall back to index.html for client side routes so refreshing /dashboard does not 404.
In detail
try_files $uri $uri/ /index.html makes deep links work in React or Vue. Hashed assets (app.3f9c1.js) get a year long immutable cache; index.html must never be cached long, or users keep loading old bundles. Proxy /api to Node from the same server block to avoid CORS entirely.
server { listen443sslhttp2; server_name example.com www.example.com; root /srv/web/current; # output of npm run build index index.html;# Hashed assets: cache forever location /assets/ { expires 1y; add_header Cache-Control "public, max-age=31536000, immutable"; try_files $uri =404; }# The HTML shell: always revalidate so new deploys show up location = /index.html { add_header Cache-Control "no-cache"; }# API on the same origin, no CORS needed location /api/ { proxy_passhttp://127.0.0.1:3000/; proxy_set_header Host $host; proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; proxy_set_header X-Forwarded-Proto $scheme; }# Client side routes fall back to the SPA location / { try_files $uri $uri/ /index.html; }}
Why it matters Caching index.html for even an hour leaves users on old JavaScript after every deploy.
Certbot gets a free certificate from Let's Encrypt, edits your Nginx config to use it, and renews it automatically every 60 days.
In detail
Run certbot --nginx -d example.com -d www.example.com, choose redirect, and confirm the systemd timer exists for renewals. Redirect all HTTP to HTTPS, enable HTTP/2, and add HSTS once you are sure HTTPS works everywhere. On AWS, certificates from ACM attach to load balancers and CloudFront for free, but cannot be installed on a plain server.
Turn on gzip (or brotli) for text, send long cache headers for hashed files, and let Nginx buffer slow clients so Node workers are freed immediately.
In detail
Compression cuts JSON and JavaScript by 70 percent or more; skip images and videos, which are already compressed. Open file caching speeds up static serving. Proxy buffering means a phone on a slow network downloads from Nginx, not from Node. For heavy, cacheable API responses, proxy_cache can store them on disk for seconds or minutes.
WebSockets start as an HTTP request that asks to upgrade. Nginx must forward the Upgrade and Connection headers and keep idle connections open long enough.
In detail
Add proxy_http_version 1.1, the Upgrade headers and a long proxy_read_timeout for the socket path. With PM2 cluster mode, Socket.io needs sticky sessions or the Redis adapter, because a client's polling requests may land on different workers. Using the websocket transport only, plus the Redis adapter, avoids stickiness.
An upstream block lists several backends; Nginx spreads requests across them with round robin, least connections or IP hash, and skips ones that fail.
In detail
Use it to balance across PM2 instances listening on different ports, or across two or three Linodes on a private network. max_fails and fail_timeout give passive health checking. A backup server only gets traffic when all primaries are down. For true active health checks and autoscaling, an AWS Application Load Balancer or Linode NodeBalancer does it as a managed service.
upstream.confnginx
upstream api { least_conn; server 10.0.0.2:3000; server 10.0.0.3:3000; }
limit_req caps requests per second per client IP at the edge, before they cost Node anything; limit_conn caps simultaneous connections; client_max_body_size caps uploads.
In detail
Define zones in the http block and apply them per location: stricter for login and OTP routes, looser for reads. burst lets short spikes through with nodelay. Return 429, and log rejections so you can tune. Edge limits complement, not replace, per user limits in the app with Redis.
Add headers that tell browsers to always use HTTPS, refuse framing, sniffing and unexpected scripts, and hide the Nginx version.
In detail
HSTS forces HTTPS for a year once a browser has seen it; only enable it when every subdomain works over HTTPS. X-Content-Type-Options stops MIME sniffing, frame-ancestors in a Content Security Policy blocks clickjacking, and Referrer-Policy limits leaked URLs. A full CSP for the frontend is the strongest XSS defence but needs testing in report only mode first. Block access to dotfiles such as .env and .git.
nginx -t checks syntax and file paths; systemctl reload nginx starts new workers with the new config while old workers finish their in flight requests. No connection is dropped.
In detail
Restart kills everything and briefly refuses connections; reload does not. If the test fails, the running Nginx keeps the old config, so a typo cannot take the site down as long as you always test first. Keep configs in git, deploy them with your app or with Ansible, and reload in the same script that switches upstreams for blue green deploys.
reload.shShell
sudo nginx -t && sudo systemctl reload nginx
#!/usr/bin/env bash# Safe Nginx config deployset -euo pipefailsudo cp nginx/app.conf /etc/nginx/sites-available/app.conf.newsudo mv /etc/nginx/sites-available/app.conf.new /etc/nginx/sites-available/app.confif sudo nginx -t; then sudo systemctl reload nginx # graceful: no dropped connections echo "nginx reloaded"else echo "config test failed, running config untouched" >&2 exit 1fi# What happens on reload:# master reads new config -> starts new workers# old workers stop accepting, finish in flight requests, exit
Why it matters Because reload only applies a config that passed the test, the worst a typo can do is fail the deploy, not the site.
The access log records every request with status, size and timing; the error log records proxy failures, permission problems and config issues. A JSON log format makes both easy to ship to CloudWatch and query.
In detail
Add request_time and upstream_response_time to the format: the difference shows whether slowness is in Nginx and the network or in Node. Include $request_id and pass it to Node so one id links Nginx, app and database logs. Logrotate keeps the files from filling the disk; Ubuntu's package already rotates daily.
node server.js dies when it crashes, when you close SSH, and when the server reboots. A process manager restarts it, runs one copy per CPU core, collects logs and reloads code without downtime.
In detail
PM2 is the common choice for Node on a VPS: it daemonises your app, restarts on crash with backoff, offers cluster mode, graceful reloads, log files with rotation and a boot script. systemd can do the restart and boot part for any program and is a fine alternative for single process apps; containers hand the job to Docker or ECS. On Lambda there is no long running process at all, so none of this applies.
pm2 start launches and names a process; status, logs, restart, reload, stop and delete manage it; describe shows everything about one app.
In detail
Name every process so scripts can target it. pm2 status shows uptime, restarts and memory; a rising restart count means crashes you should read about in pm2 logs. pm2 describe shows the exact script path, node version and environment PM2 is using, which settles most 'works on my machine' puzzles. Run PM2 as the deploy user, never root.
pm2.shShell
pm2 start dist/main.js --name api && pm2 status && pm2 logs api
pm2 start dist/main.js --name api # start and name itpm2 status # table: status, cpu, memory, restarts, uptimepm2 logs api --lines100# tail stdout and stderrpm2 describe api # script, cwd, node version, env, log pathspm2 restart api # hard restart (brief downtime)pm2 reload api # zero downtime reload (cluster mode)pm2 stop api # stop but keep in the listpm2 delete api # remove from the listpm2 monit # live terminal dashboardpm2 save # remember the current list for reboots
Why it matters Watch the restart column in pm2 status: a number that keeps climbing is a crash loop hiding behind auto restarts.
ecosystem.config.js declares every process: script, working directory, instances, environment, log paths, memory limits and graceful shutdown timings. Commit it to the repo.
In detail
Point cwd at the current symlink so reloads pick up new releases. Use env_production for production values and load secrets from a file outside the repo with dotenv or from SSM at startup. max_memory_restart restarts a worker that leaks past a threshold; listen_timeout and kill_timeout control how long PM2 waits for a worker to become ready and to shut down cleanly.
// /srv/api/ecosystem.config.jsmodule.exports = { apps: [ { name: "api", cwd: "/srv/api/current", // follows the release symlink script: "dist/main.js", instances: "max", // one worker per CPU core exec_mode: "cluster", node_args: "--enable-source-maps", max_memory_restart: "600M", // recycle a leaking worker wait_ready: true, // wait for process.send("ready") listen_timeout: 10000, kill_timeout: 8000, // time to finish in flight requests out_file: "/srv/api/shared/logs/api.out.log", error_file: "/srv/api/shared/logs/api.err.log", merge_logs: true, time: true, // timestamp every log line env_production: {NODE_ENV: "production",PORT: 3000,DOTENV_CONFIG_PATH: "/srv/api/shared/.env", }, }, { name: "worker", cwd: "/srv/api/current", script: "dist/worker.js", instances: 1, env_production: { NODE_ENV: "production" }, }, ],};// pm2 start ecosystem.config.js --env production
Why it matters wait_ready with listen_timeout is what makes reloads truly zero downtime: PM2 only kills an old worker after the new one says it is ready.
Node runs your JavaScript on one CPU core. PM2 cluster mode starts one worker per core and shares the port between them, so a four core server handles roughly four times the load.
In detail
Workers share nothing in memory, so sessions, caches and rate limits must live in Redis or the database, and in memory timers run once per worker (move cron jobs to a single worker or a queue). Leave a core free if the server also runs Redis or heavy Nginx work. Cluster mode is also what enables pm2 reload, because it can replace workers one at a time.
Back of the envelope
Workersone per coreminus one if Redis is local
Throughputabout cores times onefor CPU bound code
Memoryeach worker separatewatch total RSS
Cronworker 0 onlyvia NODE_APP_INSTANCE
cluster.shShell
pm2 start dist/main.js -i max --name api && pm2 scale api 3
pm2 start dist/main.js -i max --name api # one worker per corepm2 scale api 3# set exactly 3 workerspm2 scale api +1# add one more# Prove the load spreadsfor i in $(seq 18); do curl -s localhost:3000/whoami; echo; done# {"pid":4211} {"pid":4218} {"pid":4225} ... different workers# In the app, run scheduled jobs on one worker only# if (process.env.NODE_APP_INSTANCE === "0") startCronJobs();
Why it matters NODE_APP_INSTANCE is set by PM2 per worker, which is the simplest way to stop a cron job running four times.
pm2 reload replaces cluster workers one by one: start a new worker, wait until it is ready, tell an old worker to stop, let it finish in flight requests, repeat. Users never see a refused connection.
In detail
Your app must cooperate. On start, call process.send('ready') once the server is listening and the database is connected. On SIGINT, stop accepting new connections, finish current requests, close the database pool and exit within kill_timeout. Without this, reload falls back to killing workers mid request. Restart, unlike reload, stops everything at once and causes a brief outage.
Start, wait for ready, drain the old oneEdgeExternalService
PM2 writes each app's stdout and stderr to files. pm2-logrotate rotates them by size or date, compresses old ones and keeps a fixed number, so a chatty app never fills the disk.
In detail
Log JSON lines to stdout from the app (pino or winston), let PM2 write them to files with timestamps, rotate with pm2-logrotate, and ship the files to CloudWatch with the agent. pm2 flush empties logs, useful after an incident is resolved. Never log secrets, tokens or full card numbers.
logrotate.shShell
pm2 install pm2-logrotate && pm2 set pm2-logrotate:max_size 50M && pm2 set pm2-logrotate:retain 14
pm2 install pm2-logrotatepm2 set pm2-logrotate:max_size 50M # rotate when a file reaches 50 MBpm2 set pm2-logrotate:retain 14# keep 14 rotated filespm2 set pm2-logrotate:compress true # gzip old filespm2 set pm2-logrotate:rotateInterval '0 0 * * *'# and daily at midnightpm2 logs api --lines200 --raw | grep '"level":50'# errors only (pino level 50)pm2 flush api # empty the current filesdf -h /srv # keep an eye on disk space
Why it matters Disk full is one of the most common causes of a mysterious outage; rotation with a retain limit removes it entirely.
pm2 startup generates a systemd unit that starts PM2 at boot for your user, and pm2 save records the current process list so it is restored automatically.
In detail
Run pm2 startup as the deploy user, copy the sudo command it prints, then start your apps and run pm2 save. Repeat pm2 save after adding or removing apps. Test it with a real reboot during a quiet window; kernel updates will reboot the box eventually anyway.
startup.shShell
pm2 startup systemd -u deploy --hp /home/deploy && pm2 save
# As the deploy userpm2 startup systemd# PM2 prints a command like this; run it once:sudo env PATH=$PATH:/usr/bin pm2 startup systemd -u deploy --hp /home/deploypm2 start /srv/api/ecosystem.config.js --env productionpm2 save # snapshot the list to ~/.pm2/dump.pm2# Verify with a real reboot in a quiet windowsudo reboot# ...after it comes backpm2 statussystemctl status pm2-deploy
Why it matters Forgetting pm2 save is the classic reason an app is missing after the first reboot.
pm2 monit shows live CPU, memory and logs per worker; pm2 status shows restarts; max_memory_restart recycles leaking workers; and Node's own metrics tell you about the event loop.
In detail
Watch three numbers: restart count (crashes), memory per worker over days (leaks) and event loop lag (blocked CPU). Expose a /metrics or /health endpoint with process.memoryUsage and loop delay, and scrape it from CloudWatch or an uptime checker. A memory restart is a safety net, not a fix; take a heap snapshot to find the leak.
A deploy script builds or uploads into a new release folder, installs dependencies, runs checks, flips the current symlink, reloads PM2 and verifies health, rolling back automatically if the check fails.
In detail
Everything slow or risky happens before the switch; the switch itself is one system call. Shared files are linked into the release. After reload, poll the health endpoint for a few seconds; on failure, point current back to the previous release and reload again. Keep the last five releases and prune the rest.
Run two copies of the app on different ports (blue on 3001, green on 3002). Deploy to the idle one, test it on its port, then point the Nginx upstream at it and reload. The old colour stays warm for instant rollback.
In detail
This works even for apps that cannot use PM2 cluster reloads, such as a single process with long startup, and lets you smoke test the new version on the real server before any user sees it. It costs double memory during the switch. Across several servers the same idea becomes an AWS target group swap or a DNS weight change.
Deploy to the idle colour, then flip NginxEdgeServiceClientData
#!/usr/bin/env bash# blue = 3001, green = 3002; the active one is in /etc/nginx/conf.d/api_active.incset -euo pipefailACTIVE=$(grep -oE'[0-9]{4}' /etc/nginx/conf.d/api_active.inc)if [ "$ACTIVE" = "3001" ]; then IDLE=3002; COLOR=green; else IDLE=3001; COLOR=blue; fipm2 reload "api-$COLOR" --update-env# deploy to the idle colourfor i in $(seq 115); do curl -fsS"http://127.0.0.1:$IDLE/health" && break; sleep 2; donecurl -fsS"http://127.0.0.1:$IDLE/health" >/dev/null || { echo "idle colour unhealthy"; exit 1; }echo "server 127.0.0.1:$IDLE;" | sudo tee /etc/nginx/conf.d/api_active.incsudo nginx -t && sudo systemctl reload nginxecho "live on $COLOR ($IDLE); previous colour stays up for rollback"# nginx: upstream api { include /etc/nginx/conf.d/api_active.inc; keepalive 32; }
Why it matters Testing the idle colour on its own port before switching means users only ever meet a version that already passed its checks.
During a deploy, old and new code run at the same time for a while, so every migration must work with both. Use expand and contract: add first, migrate data, switch code, remove later.
In detail
Never rename or drop a column in the same deploy that stops using it. Add nullable columns, backfill in batches, then add constraints. Create indexes concurrently in PostgreSQL to avoid locking writes. Run migrations as a separate step before switching traffic, with a lock timeout so a long lock fails fast instead of freezing production. Test migrations against a copy of production data.
migration.sqlSQL
SET lock_timeout = '3s';ALTER TABLE users ADD COLUMN display_name text; -- expandCREATEINDEX CONCURRENTLY users_display_name_idx ON users (display_name);
-- Release 1: expand (safe with old and new code)SET lock_timeout = '3s'; -- fail fast instead of blocking writesSET statement_timeout = '15min';ALTER TABLE users ADD COLUMN display_name text; -- nullable, instantCREATEINDEX CONCURRENTLY IF NOT EXISTS users_display_name_idx ON users (display_name);-- Backfill in batches from a script, not one giant UPDATEUPDATE users SET display_name = first_name || ' ' || last_nameWHERE id IN (SELECT id FROM users WHERE display_name ISNULLLIMIT5000);-- Release 2: code reads and writes display_name only-- Release 3: contract (after release 2 is stable everywhere)ALTER TABLE users ALTER COLUMN display_name SETNOTNULL;ALTER TABLE users DROP COLUMN first_name, DROP COLUMN last_name;
Why it matters Spreading one change across three releases means any single deploy can be rolled back without touching the database.
On every push, CI installs, lints, tests and builds; on main, it packages an artifact and deploys over SSH to the server or uploads to Lambda and S3, then runs a smoke test.
In detail
Build once in CI and ship the artifact; never build on production servers. Store the SSH key and AWS credentials as repository secrets, or better, use GitHub's OIDC to assume an AWS role without long lived keys. Protect main with required checks, use environments for manual approval to production, and limit concurrency so two deploys never overlap.
Build once, deploy the same artifact everywhereClientServiceDataExternal
git pushClientLint, test, buildServiceArtifactDataLinode via SSHServiceLambda + S3Externalon mainrelease.shOIDC role
deploy.ymlYAML
- run: npm ci && npm test && npm run build- run: scp api.tgz deploy@$HOST:/tmp/ && ssh deploy@$HOST "/srv/api/release.sh /tmp/api.tgz"
Every deploy needs a way back that takes under a minute: repoint the symlink and reload, flip the blue green upstream, publish the previous Lambda version, or redeploy the previous frontend build.
In detail
Decide rollback triggers in advance: error rate above baseline, health check failures, or a key business metric dropping. Roll back first, investigate second. Database changes are the hard part, which is why expand and contract migrations matter: if the schema is compatible with the previous code, rollback is just code. Practise it, so the commands are muscle memory.
rollback.shShell
ln -sfn"$(ls -1dt /srv/api/releases/* | sed -n 2p)" /srv/api/current && pm2 reload api
#!/usr/bin/env bash# One command back to the previous releaseset -euo pipefailBASE=/srv/apiPREV=$(ls -1dt "$BASE"/releases/* | sed -n 2p) # second newestecho "rolling back to $PREV"ln -sfn"$PREV""$BASE/current.tmp" && mv -Tf "$BASE/current.tmp""$BASE/current"pm2 reload "$BASE/ecosystem.config.js" --env production --update-envcurl -fsS http://127.0.0.1:3000/health && echo "rolled back"# Lambda: point the live alias at the previous version# aws lambda update-alias --function-name orders-api --name live --function-version 41# Frontend on S3: re-sync the previous build and invalidate# aws s3 sync builds/2026-10-05/ s3://acme-web/ --delete# aws cloudfront create-invalidation --distribution-id E123 --paths "/index.html"
Why it matters Having one rollback command per platform, written before the incident, is what keeps a bad deploy to a two minute blip.
A liveness endpoint answers 'is the process running'; a readiness endpoint answers 'can it serve traffic', checking the database and other critical dependencies. Deploy scripts, load balancers and uptime monitors all poll them.
In detail
Keep /health cheap and dependency free so a slow database does not cause restarts. Make /ready check the database with a short timeout and return 503 while starting up or shutting down. Exclude health routes from auth, rate limits and access logs. On AWS, target groups and Route 53 health checks use the same endpoints.
A React, Vue or Svelte app builds into a folder of static files: one index.html, hashed JavaScript and CSS bundles and assets. That folder is the whole frontend deployment.
In detail
Build in CI with the same Node version every time, with production mode on so code is minified and dead code removed. The output is pure static files: serve them from Nginx on your server, or from S3 behind CloudFront. Check bundle size in CI; a sudden jump usually means an accidental heavy import. Source maps help debugging but should not be public, or upload them only to your error tracker.
build.shShell
npm ci && npm run build && ls dist/assets
npm cinpm run build # vite build -> dist/# dist/# index.html short cache, references hashed files# assets/index-3f9c1a2b.js long cache, name changes when content changes# assets/index-8d02e4f1.css# favicon.svgdu -sh dist && find dist -name"*.js" -size +300k # flag large bundles# Ship to your serverrsync -az --delete dist/ deploy@203.0.113.10:/srv/web/releases/$(date +%Y%m%dT%H%M%S)/
Why it matters Because every bundle name contains a content hash, an old browser tab and a new deploy never fight over the same file.
Upload the build to a private S3 bucket and put CloudFront in front: it caches files at edge locations worldwide, serves HTTPS with a free ACM certificate and routes /api to your backend if you want one domain.
In detail
Keep the bucket private and grant CloudFront access with origin access control. Upload hashed assets with a one year cache and index.html with no cache, then invalidate only /index.html after each deploy. For SPA routes, return index.html for 403 and 404 from S3. A second behaviour can forward /api/* to your Linode or an API Gateway, avoiding CORS.
One domain, edge cached files, API behind the same doorClientEdgeDataServiceExternal
BrowserClientCloudFront edgeEdgeS3 bucket (private)DataAPI on Linode or LambdaServiceACM certificateExternalHTTPS/*/api/*
Give every built file a content hash in its name and cache it forever; keep the HTML that references them uncached. Users always get the latest app, and repeat visits load almost nothing.
In detail
Cache-Control: public, max-age=31536000, immutable for /assets, and no-cache for index.html, which still allows a fast 304 Not Modified check. Service workers add another cache layer that can pin old versions if misconfigured. For API responses, send no-store on personal data and short public max-age only on shared, cacheable data.
Frontend environment variables are inlined into JavaScript at build time and anyone can read them. Only public values belong there: API base URLs, analytics keys, feature flags.
In detail
Vite exposes only variables prefixed with VITE_, Next.js only NEXT_PUBLIC_. Build once per environment, or, to build once and deploy everywhere, serve a small /config.json from the server that the app fetches at startup. Secrets such as database URLs or private API keys must stay on the server, behind your own API.
Next.js with server rendering is a Node app, not static files. Build with output: 'standalone', run the generated server with PM2, put Nginx in front for TLS and caching of /_next/static.
In detail
The standalone output contains a minimal server.js and only the node_modules it needs, so releases are small. Copy .next/static and public alongside it. Nginx serves /_next/static with a long immutable cache and proxies everything else. A fully static export (output: 'export') can go to S3 and CloudFront instead.
next.shJavaScript
next build && cp -r .next/static .next/standalone/.next/ && pm2 start .next/standalone/server.js --name web
IAM decides who can do what in AWS. Lock the root account away with MFA, give people their own users or SSO logins, and give apps roles with only the permissions they need.
In detail
Never use root for daily work or put access keys in code. EC2 instances and Lambda functions get IAM roles, so credentials rotate automatically. CI uses OIDC to assume a deploy role. Write policies that name exact actions and resources, for example s3:PutObject on one bucket prefix, instead of s3:*. Turn on CloudTrail so every API call is recorded, and set a budget alarm on day one.
An EC2 instance and a Linode are both virtual servers you SSH into and set up the same way. The difference is everything around them: VPCs, security groups, IAM roles and managed services on AWS versus simplicity and predictable pricing on Linode.
In detail
Choose Linode when you want one or a few servers, fixed monthly bills and a simple dashboard. Choose EC2 when the app leans on other AWS services (RDS, S3, Lambda, SQS) and should reach them over a private network with IAM roles instead of keys. Mixing is fine: a Linode API can use RDS over TLS, though cross provider traffic adds latency and egress cost.
RDS runs PostgreSQL for you: installation, patching, automated backups, point in time recovery, monitoring and optional Multi AZ failover. You pick the instance size, storage and network placement.
In detail
Start with a burstable db.t4g instance for small apps, gp3 storage with autoscaling, a parameter group for settings such as statement_timeout and slow query logging, and Performance Insights on. Multi AZ keeps a synchronous standby in another zone and fails over in about a minute; read replicas scale reads. Size connections: each Node worker holds a pool, so total connections equal servers times workers times pool size.
Put RDS in private subnets with public access off, allow port 5432 only from the app's security group, and require TLS. Connect from your laptop through a bastion, SSM Session Manager or a VPN, never by opening the database to the internet.
In detail
Security groups can reference other security groups, so 'allow from sg-app' keeps working as servers come and go. If the app runs on Linode outside AWS, allow only the Linode's static IP and enforce rds.force_ssl. Rotate the master password through Secrets Manager and create a separate, least privileged database user for the app.
Only the app's security group and an audited tunnel reach the databaseClientEdgeServiceData
Your laptopClientSSM tunnelEdgeApp servers (sg-app)ServiceRDS private subnetData5432 TLSport forward
# 1. App user with only what it needs (run once as the master user)# CREATE ROLE app LOGIN PASSWORD '...';# GRANT CONNECT ON DATABASE prod TO app;# GRANT USAGE ON SCHEMA public TO app;# GRANT SELECT, INSERT, UPDATE, DELETE ON ALL TABLES IN SCHEMA public TO app;# 2. Security group: Postgres only from the app servers' groupaws ec2 authorize-security-group-ingress \ --group-id sg-0db0000000000000 \ --protocol tcp --port5432 \ --source-group sg-0app000000000000# 3. Reach the private database from your laptop without opening itaws ssm start-session --target i-0bastion00000000 \ --document-name AWS-StartPortForwardingSessionToRemoteHost \ --parameters host=prod.abc123.ap-south-1.rds.amazonaws.com,portNumber=5432,localPortNumber=5433psql"host=localhost port=5433 dbname=prod user=readonly sslmode=verify-full sslrootcert=rds-global-bundle.pem"
Why it matters Session Manager tunnels need no open SSH port and leave an audit trail, which beats a public bastion host.
RDS takes daily snapshots and keeps transaction logs, so you can restore to any second within the retention window. Manual snapshots stay until you delete them; take one before risky migrations.
In detail
A restore always creates a new instance; you then point the app at it or rename. Practise the restore and time it. Copy snapshots to another region for disaster recovery, and enable deletion protection on production instances. For long term or cross provider safety, add a nightly pg_dump to S3.
Lambda runs a function in response to an event: an HTTP request through API Gateway, a file landing in S3, a message on SQS, or a schedule. You pay per request and per millisecond, and it scales to zero.
In detail
Lambda suits spiky or background work: image resizing, webhooks, scheduled reports, queue consumers. Keep handlers small, create clients and database connections outside the handler so warm invocations reuse them, and set memory (which also sets CPU) by measuring. Cold starts add a few hundred milliseconds for Node; provisioned concurrency removes them for latency critical paths. The 15 minute limit rules out long jobs.
Back of the envelope
Timeout limit15 minutesper invocation
Memory128 MB to 10 GBCPU scales with it
Node cold startabout 200 to 600 mssmaller bundles start faster
Free tier1 M requests per monthplus 400k GB seconds
Deploying Lambda with API Gateway, versions and aliases
Bundle the function with esbuild, upload it, publish an immutable version, and move a live alias to it. API Gateway invokes the alias, so rollback is moving the alias back.
In detail
Use the AWS SAM or Serverless Framework, or plain CLI commands in CI. HTTP APIs in API Gateway are cheaper and simpler than REST APIs for most cases. Weighted aliases allow canary releases: send 10 percent to the new version, watch CloudWatch errors, then shift the rest. Set reserved concurrency to protect your database from a sudden burst.
Shift traffic between immutable versionsClientEdgeService
Each Lambda instance opens its own database connection, so a traffic spike of 1,000 concurrent functions can exhaust PostgreSQL. RDS Proxy pools and shares connections, survives failovers and can use IAM authentication.
In detail
Put the Lambda in the same VPC private subnets as the proxy, give it a security group allowed by the proxy, and connect to the proxy endpoint instead of the database. Keep the pg client outside the handler with a pool size of 1. Set reserved concurrency so even the proxy is not overwhelmed. For light, bursty workloads, the RDS Data API is an alternative over HTTPS.
db-lambda.mjsJavaScript
const pool = new pg.Pool({ host: process.env.PROXY_HOST, max: 1 });exportconsthandler = async () => (await pool.query("select now()")).rows[0];
import pg from"pg";import { Signer } from"@aws-sdk/rds-signer";const signer = newSigner({ hostname: process.env.PROXY_HOST, port: 5432, username: "app", region: process.env.AWS_REGION });// Outside the handler: reused across warm invocations, one connection per instanceconst pool = new pg.Pool({ host: process.env.PROXY_HOST, // RDS Proxy endpoint, not the DB database: "prod", user: "app", password: () => signer.getAuthToken(), // IAM auth token, no stored password ssl: { rejectUnauthorized: true }, max: 1, idleTimeoutMillis: 60_000,});exportconsthandler = async (event) => {const { rows } = await pool.query("select id, status from orders where id = $1", [event.pathParameters.id]);return rows.length ? { statusCode: 200, body: JSON.stringify(rows[0]) } : { statusCode: 404, body: "not found" };};
Why it matters IAM tokens through the proxy mean no database password lives in Lambda configuration at all.
Store user files in S3, not on the server's disk, so any server or Lambda can reach them and deploys never lose them. Let browsers upload directly with pre signed URLs.
In detail
The API creates a short lived pre signed PUT URL for a specific key and content type; the browser uploads straight to S3; an S3 event triggers a Lambda for thumbnails or virus scanning. Keep buckets private with Block Public Access on, serve files through CloudFront or pre signed GET URLs, turn on versioning, and add lifecycle rules to move old files to cheaper storage.
Keep database passwords, API keys and signing secrets in AWS Secrets Manager or SSM Parameter Store, read them at startup with the instance or Lambda role, and never commit them or bake them into images.
In detail
Parameter Store SecureString is free for standard parameters and fine for most config; Secrets Manager costs a little but adds automatic rotation, notably for RDS passwords. On a Linode outside AWS, a root owned .env with permissions 600 in the shared folder is acceptable, or fetch from AWS with a tightly scoped access key. Rotate anything that was ever pasted into chat or a ticket.
Set an AWS Budget with email alerts the day you create the account, tag resources by project, and review Cost Explorer monthly. Most surprise bills come from forgotten instances, NAT gateways, data transfer and unlimited log retention.
In detail
Right size instances after watching real usage, use Graviton (t4g, m7g) for cheaper compute, set CloudWatch log retention instead of forever, move old S3 objects to infrequent access, and delete unattached volumes and old snapshots. A NAT gateway costs money every hour plus per gigabyte; VPC endpoints for S3 avoid some of it. Linode's flat pricing is easier to predict for steady workloads.
Log one JSON object per line with a level, message, request id and useful fields. JSON logs can be searched, filtered and graphed in CloudWatch Logs Insights; plain sentences cannot.
In detail
Use pino for speed. Attach a request id (from Nginx's X-Request-Id) to every line via a child logger, log the start and end of each request with status and duration, and log errors with their stack. Redact authorization headers, passwords and tokens at the logger level. Use levels consistently: info for business events, warn for recoverable problems, error for failures that need a look.
Shipping logs and metrics with the CloudWatch agent
The CloudWatch agent tails Nginx and PM2 log files and sends them to CloudWatch Logs, and also reports memory and disk metrics that EC2 does not collect by default. It works on Linode too, with an IAM user key.
In detail
Create one log group per source (/app/api, /nginx/access, /nginx/error), set retention, and use the instance id or hostname as the stream name. On EC2 attach the CloudWatchAgentServerPolicy to the instance role; on Linode use a dedicated IAM user limited to those log groups. Lambda sends its console output to CloudWatch automatically.
Servers and functions land in one placeEdgeServiceExternal
Logs Insights queries JSON logs across groups with a small pipe language: filter, stats, sort. Find the slowest routes, count 5xx by path, or follow one request id across Nginx and the app.
In detail
Save the handful of queries you reach for during incidents. Because logs are JSON, fields such as status, rt and route are queryable without parsing. Queries are billed by data scanned, so narrow the time range first. Metric filters can turn a log pattern (such as level error) into a CloudWatch metric for alarms.
queries.txtNotes
fields @timestamp, uri, status, rt | filter status >= 500 | stats count() by uri | sort count() desc
# Slowest endpoints in the last hour (Nginx JSON log group /nginx/access)fields uri, rt| stats avg(rt) as avg_s, pct(rt, 99) as p99_s, count() as hits by uri| sort p99_s desc| limit 20# 5xx by pathfilter status >= 500| stats count() as errors by uri, status| sort errors desc# Follow one request across Nginx and the app (select both log groups)filter request_id = "8c1f0e2b7a" or req.id = "8c1f0e2b7a"| sort @timestamp asc# App errors with messages (/app/api)filter level >= 50| fields @timestamp, msg, err.message, req.url| sort @timestamp desc| limit 50
Why it matters Querying p99 rather than average shows the slow tail that real users complain about.
Alarms watch a metric and notify you through SNS by email, SMS or a chat webhook when it crosses a threshold for a few minutes: high 5xx rate, RDS CPU or free storage, Lambda errors, disk usage, or a missing heartbeat.
In detail
Alert on symptoms users feel (errors, latency, availability) more than on causes (CPU). Use several evaluation periods to avoid flapping and treat missing data as breaching for heartbeats. Start with a short list and tune it: an alarm that fires weekly for nothing trains people to ignore it.
An external uptime check requests your site from several locations every minute and alerts if it fails or gets slow. It catches what internal monitoring misses: DNS mistakes, expired certificates, firewall changes, a dead server.
In detail
Check the homepage, the API /ready endpoint and one critical flow. Route 53 health checks, UptimeRobot, Better Stack or Checkly all work. Monitor certificate expiry too. Point alerts at a channel humans actually watch, and give each alert a link to the runbook.
# A Route 53 health check on the readiness endpoint, from many regionsaws route53 create-health-check --caller-reference"api-ready-$(date +%s)" \ --health-check-config '{"Type": "HTTPS", "FullyQualifiedDomainName": "api.example.com", "Port": 443,"ResourcePath": "/ready", "RequestInterval": 30, "FailureThreshold": 3, "EnableSNI": true }'# Certificate expiry check you can run from cron anywhereEND=$(echo | openssl s_client -servername example.com -connect example.com:4432>/dev/null | openssl x509 -noout -enddate | cut -d= -f2)DAYS=$(( ( $(date -d"$END" +%s) - $(date +%s) ) / 86400 ))[ "$DAYS" -lt14 ] && echo "certificate expires in $DAYS days" | mail -s"cert warning" ops@example.com
Why it matters Checking from outside every 30 seconds in several regions catches the outages that every internal metric would show as perfectly normal.
When production misbehaves, check in order: is the process up (pm2 status), is Nginx happy (error log), are requests failing (access log status counts), is the box healthy (CPU, memory, disk), and is the database slow or out of connections.
In detail
Use read only commands first and change one thing at a time. ss -tlnp shows what listens on which port, htop shows CPU and memory per process, df -h and du find a full disk, journalctl shows system logs, and pg_stat_activity shows long running or blocked queries. Write down what you saw and did with timestamps; it becomes the incident timeline.
# 60 second triage on the serveruptime # load average vs number of coresfree -m# memory and swapdf -h / /srv # disk full?pm2 status # online? restart counts climbing?pm2 logs api --lines50 --err# recent app errorssudo ss -tlnp | grep -E ':80|:443|:3000'# what is listeningsudo tail -n50 /var/log/nginx/error.log # upstream refused or timed out?sudo journalctl -u nginx --since"15 min ago" --no-pager | tail# Database: long running and blocked queriespsql"$DATABASE_URL" -c "select pid, now() - query_start as age, state, wait_event_type, left(query, 80) from pg_stat_activity where state <> 'idle' order by age desc limit 10;"
Why it matters Running the same short list every time stops panic from skipping the obvious cause, which is usually a full disk or a crashed process.
A runbook is a short page per alert: what it means, how to confirm it, the first safe actions, how to roll back, and who to call. A blameless postmortem afterwards turns each incident into a fix.
In detail
Store runbooks next to the code and link them from every alarm. Write for someone tired at 3 a.m.: exact commands, expected output, and when to escalate. Afterwards record the timeline, impact, root cause, what went well and action items with owners. Most action items are small: a missing alarm, a better health check, a migration rule.
runbook-api-5xx.mdNotes
# API 5xx spike1. Check pm2 status and pm2 logs api --err2. If last deploy < 30 min ago: ./rollback.sh
# Runbook: API 5xx spike (alarm api-errors-spike)## Confirm (2 minutes)- CloudWatch Logs Insights: saved query "5xx by path", last 15 minutes- On a server: pm2 status (restarts climbing?), pm2 logs api --err --lines 50## Safe first actions1. Deployed in the last 30 minutes? -> /srv/api/rollback.sh, then confirm /ready2. Errors mention database timeouts? -> RDS console: CPU, connections, Performance Insights3. Disk full? -> df -h, then pm2 flush, check logrotate4. One worker stuck? -> pm2 reload api (zero downtime)## Escalate- No improvement in 15 minutes, or payments affected: page the backend lead## Afterwards- Postmortem within 2 days: timeline, impact, root cause, action items with owners
Why it matters Putting rollback as step one for recent deploys removes the most common incident in under a minute, before any deep debugging.
09
Phase 09, modules 60 to 62
Ship it like you mean it
Hardening and the full reference deployment
Patching, a production checklist and one complete deployment that uses every module.
Turn on unattended security upgrades, apply other updates weekly in a quiet window, and reboot when the kernel requires it. PM2 startup and Nginx's systemd unit bring everything back.
In detail
Ubuntu's unattended-upgrades installs security patches daily; enable automatic reboots at a fixed early morning time, or reboot manually after checking /var/run/reboot-required. With two servers behind a load balancer, update one at a time. Keep Node updated to the latest patch of your LTS line and run npm audit in CI. Snapshots before major upgrades make them reversible.
Before real users arrive, walk through security, reliability, observability and recovery. Each item links to the module that explains it, and none takes more than an hour.
In detail
Most launch day problems are small omissions: no HTTPS redirect, a debug log level, an open database port, no alarm on disk space, no tested backup, a cron job running once per worker. Tick through the list on staging first, then production, and repeat it after any big infrastructure change.
One complete, realistic setup: React on S3 and CloudFront, a Node API on two Linodes or EC2 instances behind Nginx with PM2 in cluster mode, RDS PostgreSQL in a private subnet, S3 for uploads, Lambda for image processing, secrets in AWS, logs and alarms in CloudWatch, and deploys from GitHub Actions.
In detail
Walk one request through it: DNS to CloudFront for the app shell, then /api to Nginx, which proxies to a PM2 worker, which queries RDS through a pooled TLS connection and logs JSON with the request id. An avatar upload goes straight to S3 with a pre signed URL and a Lambda makes the thumbnail. A deploy builds once in CI, releases atomically with a health checked symlink switch, and can roll back in one command.
Every module, one diagramClientEdgeServiceDataExternal
# Production deployment## Map- example.com, www -> CloudFront -> S3 (acme-web), /api/* -> api origin- api.example.com -> Nginx (2 servers) -> PM2 cluster "api" on 127.0.0.1:3000- Database -> RDS PostgreSQL 16, private subnets, Multi AZ, sg allows only sg-app- Uploads -> S3 acme-uploads (private), Lambda resize on raw/ -> thumbs/- Secrets -> Secrets Manager prod/api/*, SSM /prod/api/*- Logs -> CloudWatch /app/api, /nginx/access, /nginx/error (agent)- Alarms -> api-errors-spike, rds-prod-cpu-high, disk-used-80, api-ready-check## Deploy- API: GitHub Actions -> artifact -> release.sh (health checked, auto rollback)- Web: GitHub Actions -> s3 sync (immutable assets) -> invalidate /index.html- DB: expand and contract migrations, run before switching traffic## Recover- API rollback: /srv/api/rollback.sh- Web rollback: s3 sync builds/<sha>/ then invalidate- DB restore: point in time restore to a new instance (runbook: db-restore.md)
Why it matters A one page map like this is the document every new teammate and every 3 a.m. incident responder reads first.
Questions people ask
Short answers to the questions that come up most when deploying a Node app, each linked to the module that covers it in depth.
Do I need Nginx if Node can listen on port 443 itself?
You can, but Nginx handles TLS, compression, static files, slow clients, rate limits and graceful config reloads far better, and it lets PM2 reload Node workers behind it without dropping connections.
systemd is enough for a single process app. PM2 adds cluster mode, zero downtime reloads and log rotation for Node, which is why most Node VPS setups use it, often started by systemd at boot.
What is the difference between pm2 restart and pm2 reload?
restart stops all workers and starts them again, a brief outage. reload replaces cluster workers one at a time and waits for each to be ready, so no request is dropped.
Nginx could not reach Node: the app crashed, is still starting, listens on a different port, or is bound to another interface. Check pm2 status, pm2 logs and the Nginx error log.
Linode for a few servers with predictable bills and a simple setup. AWS when you rely on RDS, S3, Lambda and IAM roles in one private network. Mixing them is fine.
For spiky or background work: webhooks, image processing, scheduled jobs and queue consumers. Keep the main API with long lived connections and sockets on a server.
In AWS Secrets Manager or SSM Parameter Store read through IAM roles, or in a root owned .env file outside the repo on a plain VPS. Never in git, images or frontend bundles.
Build an artifact in CI, unpack it into a new release folder, install dependencies, switch the symlink, run pm2 reload with graceful shutdown, health check, and roll back automatically if it fails.
Every source linked from the modules above, grouped by the phase that uses it and then by where it lives. 174 links in total, all opening in a new tab.