Deploying the deployment roadmap
Resolving 63 modules across 10 phases
0%
Deployment roadmapBackend Engineering Guides
0 of 63 done
Focus mode on
00 to 62Ship it
Keep it up
Sleep well

Deploy it,from laptop to always on

Take a Node API and its frontend from your machine to production: a Linode you can trust, Nginx at the front door, PM2 reloading without dropping a request, releases you can roll back in a minute, and the AWS pieces you actually use, RDS, Lambda, S3 and CloudWatch, with logs and alarms that tell you before users do.

Best for: Node and React apps, Solo developers, Small backend teams

Phase 00, module 00

Before you start

The deployment dictionary

Every word this guide uses, explained in plain words, in technical terms, and with an everyday comparison.

At a glance

Node and React apps63 modules10 phases53 dictionary wordsKeyboard first ?

The stack you will wire

NginxPM2LambdaRDSCloudWatchLinode

Best way to use it

  • New to deploying

    Start at module 01 and go in order. Each module leans on the one before it.

  • Already run servers

    Open the menu with M and jump straight to the phase you need.

  • Track it

    Press D to mark a module done. Progress stays in this browser.

  • Focus on one

    Press O to read one module alone, then ← → to move between them.

Module 00 of 62, phase 00

A word for everything

The deployment dictionary

A catalog of every term in this guide: what it means in plain words, what it means technically, and an everyday comparison, each linked to the module that teaches it.

In detail

Filter by kind or type a word to find it. Terms such as reverse proxy, upstream, cluster mode or security group are also linked from the code examples, so you can always jump from a config line to its meaning.

Deployment

Basics
In plain words
Putting your app where users can reach it.
Technically
Building an artifact, shipping it to a runtime, configuring it and switching traffic to it.
Learn it in module 01What deployment really means

Artifact

Basics
In plain words
The finished files you ship.
Technically
Immutable build output (dist folder, zip, image) produced once and deployed many times.
Learn it in module 01What deployment really means

Environment variable

Basics
In plain words
A setting passed to the app from outside the code.
Technically
Key value pairs in the process environment, read via process.env, used for config and secrets.
Learn it in module 04Environments, config and secrets

DNS record

Basics
In plain words
An entry that points a name to a server.
Technically
A, AAAA, CNAME, TXT, CAA records with a TTL controlling cache time.
Learn it in module 05Domains and DNS records

TTL

Basics
In plain words
How long others remember an answer.
Technically
Seconds a resolver may cache a DNS record before asking again.
Learn it in module 05Domains and DNS records

VPS

Servers
In plain words
A rented computer in a data centre.
Technically
A virtual private server with dedicated vCPU, RAM and disk on shared hardware, such as a Linode or EC2 instance.
Learn it in module 06Creating a Linode server

SSH key

Servers
In plain words
A digital key for logging into a server.
Technically
An asymmetric key pair; the server holds the public key, you keep the private key.
Learn it in module 07SSH hardening and a deploy user

Reverse proxy

Nginx
In plain words
A server that receives requests and passes them to your app.
Technically
An intermediary terminating client connections and forwarding to upstream servers with added headers.
Learn it in module 13Reverse proxy to Node

Server block

Nginx
In plain words
The settings for one website in Nginx.
Technically
A server context matched by listen and server_name, containing location blocks.
Learn it in module 12Nginx install and config structure

Upstream

Nginx
In plain words
The group of app servers Nginx sends traffic to.
Technically
A named pool of backend addresses with balancing, keepalive and failure settings.
Learn it in module 18Upstreams and load balancing

TLS certificate

Nginx
In plain words
The padlock that proves your site is really yours.
Technically
An X.509 certificate signed by a CA, enabling HTTPS; Let's Encrypt issues 90 day ones.
Learn it in module 15HTTPS with Let's Encrypt and certbot

try_files

Nginx
In plain words
Look for a file, and if missing, show the app's homepage.
Technically
An Nginx directive that checks paths in order and falls back to a final URI, enabling SPA routing.
Learn it in module 14Static frontend and SPA routing

Graceful reload

Nginx
In plain words
Apply new settings without dropping anyone.
Technically
Start new workers with the new config while old workers finish in flight requests.
Learn it in module 21Testing and gracefully reloading Nginx

Rate limit

Nginx
In plain words
A cap on how often someone can call you.
Technically
limit_req with a leaky bucket per key, returning 429 beyond rate plus burst.
Learn it in module 19Rate limiting and request limits

Process manager

PM2
In plain words
A program that keeps your app running.
Technically
A supervisor that starts, restarts, clusters and logs long running processes.
Learn it in module 23Why a process manager

Ecosystem file

PM2
In plain words
PM2's settings file you keep in git.
Technically
ecosystem.config.js declaring apps, instances, env, logs and timeouts.
Learn it in module 25The PM2 ecosystem file

Cluster mode

PM2
In plain words
Running one copy of the app per CPU core.
Technically
PM2 uses Node's cluster module so workers share one port and load.
Learn it in module 26Cluster mode

Log rotation

PM2
In plain words
Starting fresh log files and deleting old ones.
Technically
Rotating by size or time, compressing, and keeping a fixed number of files.
Learn it in module 28PM2 logs and rotation

Blue green

Releases
In plain words
Two copies of the app, swap which one is live.
Technically
Deploy to the idle environment, test it, then repoint traffic; keep the old one for rollback.
Learn it in module 32Blue green on a single server

Expand and contract

Releases
In plain words
Changing the database in safe small steps.
Technically
Add new schema first, migrate, switch code, then remove old schema in later releases.
Learn it in module 33Safe database migrations

CI/CD

Releases
In plain words
Automatic testing and deploying when you push code.
Technically
Continuous integration builds and tests; continuous delivery ships artifacts through gated stages.
Learn it in module 34CI/CD with GitHub Actions

Rollback

Releases
In plain words
Going back to the previous working version.
Technically
Repointing traffic to the last known good release, version or build.
Learn it in module 35Rollback strategy

Health check

Releases
In plain words
A URL that says whether the app is OK.
Technically
Liveness and readiness endpoints polled by deploy scripts, load balancers and monitors.
Learn it in module 36Health and readiness endpoints

Cache busting

Frontend
In plain words
Giving changed files new names so browsers fetch them.
Technically
Content hashes in filenames with immutable caching, plus an uncached HTML shell.
Learn it in module 39Cache busting and cache headers

Invalidation

Frontend
In plain words
Telling the CDN to forget an old copy.
Technically
A CloudFront request that expires cached objects for given paths.
Learn it in module 38S3 and CloudFront static hosting

RDS

AWS
In plain words
A managed database run by AWS.
Technically
Amazon Relational Database Service: provisioned PostgreSQL with backups, patching and Multi AZ.
Learn it in module 44Amazon RDS for PostgreSQL

Multi AZ

AWS
In plain words
A standby copy of the database in another building.
Technically
Synchronous replication to a standby in another availability zone with automatic failover.
Learn it in module 44Amazon RDS for PostgreSQL

Lambda

AWS
In plain words
Code that runs only when something calls it.
Technically
Serverless functions invoked by events, billed per request and duration.
Learn it in module 47AWS Lambda for Node

Cold start

AWS
In plain words
The small delay when a function wakes up.
Technically
Time to create a new execution environment and load code before the first invocation.
Learn it in module 47AWS Lambda for Node

RDS Proxy

AWS
In plain words
A connection sharer in front of the database.
Technically
A managed pooler multiplexing many client connections onto few database connections.
Learn it in module 49Lambda with RDS through RDS Proxy

Pre signed URL

AWS
In plain words
A temporary link that lets someone upload or download one file.
Technically
A URL signed with credentials, valid for a set time and a specific operation.
Learn it in module 50S3 for user uploads and assets

Structured log

Observability
In plain words
Logs written as data, not sentences.
Technically
One JSON object per line with level, message and fields such as request id.
Learn it in module 53Structured JSON logging in Node

Request id

Observability
In plain words
A tag that follows one request everywhere.
Technically
A unique id set at the edge, passed in headers and attached to every log line.
Learn it in module 53Structured JSON logging in Node

Logs Insights

Observability
In plain words
A search box for all your logs.
Technically
A query language with filter, stats and sort over CloudWatch log groups.
Learn it in module 55CloudWatch Logs Insights queries

Phase 01, modules 01 to 05

From laptop to the internet

Deployment foundations

What deploying really means, how a request finds your app, and where your app could live.

Module 01 of 62, phase 01

Moving out of your laptop

What deployment really means

Deploying means taking code that runs on your machine and running it on a computer that is always on, reachable by name, secure, observable and easy to update without breaking users.

In detail

A deployment has five parts: a build (compile TypeScript, bundle the frontend), an artifact (the files you ship), a runtime (Node behind a process manager, or a Lambda), configuration (environment variables and secrets, never in git) and a release process (copy, switch, verify, roll back). Production differs from development in three ways that bite: no hot reload, real secrets, and many users at once. Every module after this one hardens one of those parts.

Build once, ship the artifact, switch trafficClientServiceData
Your laptopClientCI buildServiceArtifactDataServer or LambdaServiceUsersClientgit pushbuildship + switchrequests
deploy.sh
Shell
npm ci && npm run build          # buildrsync -a dist/ server:/srv/app/   # shipssh server "pm2 reload app"       # release
Why it matters Every later module hardens one of these four steps: build, ship, switch and verify.

Module 02 of 62, phase 01

From a click to your code

How a request reaches your app

A visitor types a domain; DNS returns an IP; the browser connects over TLS to Nginx on port 443; Nginx serves the frontend files itself or forwards API calls to Node on localhost:3000, managed by PM2; Node talks to RDS and logs everything.

In detail

Knowing each hop tells you where to look when something breaks. DNS problems show as 'site not found', TLS problems as certificate warnings, Nginx problems as 502 or 504, app problems as 500s in PM2 logs, and database problems as slow responses or connection errors. In AWS the same path may start with CloudFront or an Application Load Balancer, and an API route may land on Lambda instead of a server.

Every hop is a place to look when it breaksClientExternalEdgeDataService
BrowserClientDNSExternalNginx :443EdgeStatic filesDataNode via PM2ServiceRDS PostgreSQLDataCloudWatchExternallookupHTTPS//apiSQLlogs
trace.sh
Shell
dig +short app.example.com          # DNScurl -vI https://app.example.com    # TLS + Nginx + app headers
Why it matters Each command isolates one hop, so you find the broken layer in minutes instead of guessing.

Module 03 of 62, phase 01

Rent a flat, a hotel room or a taxi

VPS, PaaS, containers and serverless

You can run an app on a server you manage (Linode, EC2), a platform that manages it for you (Elastic Beanstalk, Render), containers (ECS, Kubernetes) or serverless functions (Lambda). Each trades control for convenience and cost shape.

In detail

A VPS is cheap, predictable and teaches you everything, but patching and uptime are on you. PaaS hides servers but limits customisation. Containers give portable, reproducible runtimes at the cost of orchestration. Serverless bills per request and scales to zero, ideal for spiky APIs, cron jobs and event handlers, but has cold starts, time limits and different connection patterns for databases. Many real systems mix them: a VPS or EC2 for the main API behind Nginx, Lambda for background jobs, RDS for data.

Where a Node app can live
OptionYou manageCost shapeGreat forWatch out for
VPS: Linode, EC2OS, Nginx, PM2, patchingFlat monthlyAPIs, sockets, full controlPatching and uptime are yours
PaaS: Render, BeanstalkCode and configPer instance, higherSmall teams, fast startLess control, platform limits
Containers: ECS, KubernetesImages and orchestrationPer node or taskMany services, reproducible buildsOperational complexity
Serverless: LambdaFunctions onlyPer request, scales to zeroSpiky APIs, jobs, eventsCold starts, 15 min limit, DB connections
Static: S3 + CloudFrontBuild outputPenniesFrontend apps and sitesNo server side code

Module 04 of 62, phase 01

Same code, different settings

Environments, config and secrets

Keep code identical across development, staging and production, and change behaviour only through environment variables. Secrets such as database passwords never go in git.

In detail

The Twelve Factor rule: config lives in the environment. Locally use a .env file ignored by git; on a server load variables through PM2's ecosystem file or a systemd EnvironmentFile readable only by the app user; on AWS read them from SSM Parameter Store or Secrets Manager. Validate config at startup and refuse to boot if something is missing. Frontend builds bake public variables in at build time, so never put secrets in them.

config.js
JavaScript
const required = ["DATABASE_URL", "JWT_SECRET"];for (const k of required) if (!process.env[k]) throw new Error(`Missing ${k}`);
Why it matters Failing fast on a missing variable turns a confusing runtime bug into a one line startup error.

Module 05 of 62, phase 01

Giving your server a name

Domains and DNS records

Buy a domain, then create DNS records: an A record points a name to your server's IPv4, AAAA to IPv6, CNAME aliases one name to another, and TXT proves ownership for certificates and email.

In detail

Typical setup: example.com and www as A records to your Linode IP, api.example.com to the same server or an AWS load balancer, and the apex on AWS using an alias record to CloudFront. Lower the TTL to 300 seconds a day before a migration so changes spread fast, then raise it again. Route 53, Linode DNS Manager or Cloudflare all work; keep DNS somewhere separate from your server so a dead server never takes your DNS with it.

zone.txt
Notes
example.com.      300  A      203.0.113.10www.example.com.  300  CNAME  example.com.
Why it matters A CAA record is a one line security win: no other certificate authority can issue a certificate for your domain.

Phase 02, modules 06 to 11

Your first server

A Linode VPS from zero

Create a server, lock it down, install Node and lay out your app so deploys stay boring.

Module 06 of 62, phase 02

Renting your first computer

Creating a Linode server

Pick a region near your users, an Ubuntu LTS image and a shared CPU plan, add your SSH public key, and you have a server with a public IP in about a minute.

In detail

A 2 GB shared Linode comfortably runs Nginx, a Node API in PM2 cluster mode with two workers and a small Redis. Use a dedicated CPU plan when CPU stays above about 50 percent. Enable Linode Backups, attach a Cloud Firewall, and add a private IP if you will connect to another Linode. The same steps map to an EC2 instance on AWS: AMI instead of image, security group instead of Cloud Firewall, Elastic IP for a fixed address.

Back of the envelope

  • Plan2 GB sharedNginx + Node x2 + Redis
  • Boot timeabout 1 minuteready for SSH
  • Backupsdaily + weeklyLinode Backups add on
  • Upgrade whenCPU over 50%for hours
create.sh
Shell
linode-cli linodes create --region ap-south --type g6-standard-1 \  --image linode/ubuntu24.04 --authorized_keys "$(cat ~/.ssh/id_ed25519.pub)"
Why it matters Passing your SSH key at creation means password login is never needed, even for the first connection.

Module 07 of 62, phase 02

Only your key opens the door

SSH hardening and a deploy user

Create a non root user with sudo, log in with SSH keys only, disable root login and password authentication, and add fail2ban to block brute force attempts.

In detail

Bots start guessing passwords within minutes of a server going online. Keys make guessing pointless. A separate deploy user owns the app directory and runs PM2, so a compromised app cannot touch system files. Keep a second key or the provider's console (Linode Lish, EC2 Instance Connect) as a way back in before you restart sshd with new settings.

harden.sh
Shell
adduser deploy && usermod -aG sudo deploysed -i 's/^#\?PermitRootLogin.*/PermitRootLogin no/' /etc/ssh/sshd_config
Why it matters sshd -t validates the new config first, so a typo cannot lock you out of your own server.

Module 08 of 62, phase 02

Close every window you are not using

Firewalls: ufw, Cloud Firewall and security groups

Allow only SSH, HTTP and HTTPS from the internet. Node, Redis and PostgreSQL ports stay closed to the world and listen on localhost or a private network.

In detail

Use two layers. A provider firewall (Linode Cloud Firewall, AWS security group) drops traffic before it reaches the server. A host firewall (ufw) protects the box even if the provider rule is wrong. Restrict SSH to your office or VPN IP when you can. On AWS, security groups reference each other: the RDS group allows port 5432 only from the app servers' group, never from 0.0.0.0/0.

Two layers outside, loopback insideClientEdgeServiceData
InternetClientCloud firewallEdgeufw on hostEdgeNginx 80/443ServiceNode 127.0.0.1:3000ServicePostgres privateData22, 80, 443allowedloopback
ufw.sh
Shell
ufw default deny incoming && ufw allow OpenSSH && ufw allow "Nginx Full" && ufw enable
Why it matters Binding Node to 127.0.0.1 means even a firewall mistake cannot expose it; only Nginx can reach it.

Module 09 of 62, phase 02

The right Node, every time

Installing Node, build tools and Git

Install the current Node LTS from NodeSource or with nvm for the deploy user, plus Git and build essentials for native modules. Pin the version so every server and CI run the same one.

In detail

Distribution packages are often years behind. NodeSource gives a system wide Node managed by apt; nvm gives per user versions and is handy when several apps need different ones, but then PM2 must be started from that user's shell so it finds the same node. Write the version in .nvmrc and in package.json engines, and make CI use it too.

install-node.sh
Shell
curl -fsSL https://deb.nodesource.com/setup_22.x | sudo -E bash -sudo apt-get install -y nodejs git build-essential
Why it matters Pinning Node in .nvmrc and engines catches a mismatched runtime before it becomes a production only bug.

Module 10 of 62, phase 02

A shelf for every version

Release folders and atomic symlink switches

Deploy each version into its own timestamped folder under releases/, keep shared files such as .env and uploads in shared/, and point a current symlink at the active release. Switching the symlink is instant and rolling back is just pointing it back.

In detail

Never git pull into the live folder: half updated files serve broken pages mid deploy. With releases, the new version is fully installed and tested before it becomes current. Keep the last five releases for quick rollbacks and delete older ones. PM2 and Nginx both point at /srv/app/current, so neither config changes between releases.

layout.txt
Notes
/srv/api/releases/20261006T1015/   /srv/api/shared/.env   /srv/api/current -> releases/20261006T1015
Why it matters mv -T replaces the symlink in a single system call, so there is never a moment with no current release.

Module 11 of 62, phase 02

Assume the disk will die

Backups, snapshots and restore drills

Back up three things: the server (Linode Backups or EC2 snapshots), the database (RDS automated backups plus manual snapshots before big changes) and user files (S3 with versioning). A backup is only real once you have restored it.

In detail

Follow 3 2 1: three copies, two kinds of storage, one off site. Server images let you rebuild fast but the database is what really matters. RDS keeps point in time recovery for up to 35 days. Write the restore steps down and rehearse them every quarter, timing how long it takes; that time is your real recovery time objective.

backup.sh
Shell
pg_dump "$DATABASE_URL" -Fc -f /tmp/app.dump && aws s3 cp /tmp/app.dump s3://acme-backups/db/
Why it matters Restoring into a scratch database on a schedule is the only proof that your backups actually work.

Phase 03, modules 12 to 22

The front door

Nginx as reverse proxy and web server

Proxy to Node, serve the frontend, add HTTPS, compress, cache, rate limit and reload without dropping a request.

Module 12 of 62, phase 03

Meet the receptionist

Nginx install and config structure

Nginx is a fast web server and reverse proxy. It terminates HTTPS, serves static files, forwards API calls to Node and protects it from slow clients and bursts.

In detail

Config lives in /etc/nginx: nginx.conf holds global settings, each site gets a file in sites-available linked into sites-enabled (or a file in conf.d). Blocks nest: http contains server blocks (one per domain), which contain location blocks (one per path). Always run nginx -t before reloading. Reload is graceful: old workers finish their requests while new ones take over.

nginx.conf
nginx
http {  server { listen 80; server_name example.com; location / { root /srv/web/current; } }}
Why it matters worker_processes auto and a test before every reload are the two habits that prevent most Nginx outages.

Module 13 of 62, phase 03

Forwarding calls to the right desk

Reverse proxy to Node

A location block with proxy_pass forwards API requests to Node on localhost. Pass the original host, client IP and protocol in headers so the app can log and build URLs correctly.

In detail

Set Host, X-Real-IP, X-Forwarded-For and X-Forwarded-Proto, and tell Express to trust the proxy so req.ip and req.protocol are right. Keep connections to Node alive with an upstream keepalive pool. Tune timeouts: proxy_read_timeout of 60 seconds suits normal APIs, longer for slow exports. 502 means Node refused or crashed; 504 means it was too slow.

TLS ends at Nginx; Node speaks plain HTTP on loopbackClientEdgeService
ClientClientNginxEdgeNode worker 1ServiceNode worker 2ServiceHTTPSHTTP keepalive
api.conf
nginx
location /api/ {  proxy_pass http://127.0.0.1:3000;  proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;}
Why it matters Without trust proxy, every request looks like it came from 127.0.0.1, which breaks rate limits and logs.

Module 14 of 62, phase 03

Serving the frontend yourself

Static frontend and SPA routing

Point Nginx's root at the built frontend folder, serve files directly, and fall back to index.html for client side routes so refreshing /dashboard does not 404.

In detail

try_files $uri $uri/ /index.html makes deep links work in React or Vue. Hashed assets (app.3f9c1.js) get a year long immutable cache; index.html must never be cached long, or users keep loading old bundles. Proxy /api to Node from the same server block to avoid CORS entirely.

web.conf
nginx
root /srv/web/current;location / { try_files $uri $uri/ /index.html; }
Why it matters Caching index.html for even an hour leaves users on old JavaScript after every deploy.

Module 15 of 62, phase 03

The padlock in the address bar

HTTPS with Let's Encrypt and certbot

Certbot gets a free certificate from Let's Encrypt, edits your Nginx config to use it, and renews it automatically every 60 days.

In detail

Run certbot --nginx -d example.com -d www.example.com, choose redirect, and confirm the systemd timer exists for renewals. Redirect all HTTP to HTTPS, enable HTTP/2, and add HSTS once you are sure HTTPS works everywhere. On AWS, certificates from ACM attach to load balancers and CloudFront for free, but cannot be installed on a plain server.

certbot.sh
Shell
sudo certbot --nginx -d example.com -d www.example.com --redirect
Why it matters renew --dry-run catches a broken renewal now, instead of an expired certificate in two months.

Module 16 of 62, phase 03

Smaller and closer is faster

Compression, caching and buffers

Turn on gzip (or brotli) for text, send long cache headers for hashed files, and let Nginx buffer slow clients so Node workers are freed immediately.

In detail

Compression cuts JSON and JavaScript by 70 percent or more; skip images and videos, which are already compressed. Open file caching speeds up static serving. Proxy buffering means a phone on a slow network downloads from Nginx, not from Node. For heavy, cacheable API responses, proxy_cache can store them on disk for seconds or minutes.

perf.conf
nginx
gzip on; gzip_types application/json text/css application/javascript;
Why it matters A 30 second micro cache turns a thousand identical requests per second into two requests a minute for Node.

Module 17 of 62, phase 03

Keeping the line open

Proxying WebSockets and Socket.io

WebSockets start as an HTTP request that asks to upgrade. Nginx must forward the Upgrade and Connection headers and keep idle connections open long enough.

In detail

Add proxy_http_version 1.1, the Upgrade headers and a long proxy_read_timeout for the socket path. With PM2 cluster mode, Socket.io needs sticky sessions or the Redis adapter, because a client's polling requests may land on different workers. Using the websocket transport only, plus the Redis adapter, avoids stickiness.

ws.conf
nginx
proxy_set_header Upgrade $http_upgrade;proxy_set_header Connection "upgrade";
Why it matters The map block sends Connection: close for normal requests and upgrade for sockets from one shared location.

Module 18 of 62, phase 03

More than one kitchen

Upstreams and load balancing

An upstream block lists several backends; Nginx spreads requests across them with round robin, least connections or IP hash, and skips ones that fail.

In detail

Use it to balance across PM2 instances listening on different ports, or across two or three Linodes on a private network. max_fails and fail_timeout give passive health checking. A backup server only gets traffic when all primaries are down. For true active health checks and autoscaling, an AWS Application Load Balancer or Linode NodeBalancer does it as a managed service.

upstream.conf
nginx
upstream api { least_conn; server 10.0.0.2:3000; server 10.0.0.3:3000; }
Why it matters proxy_next_upstream retries a failed request on another server, so one crashing box costs users a few milliseconds, not an error.

Module 19 of 62, phase 03

Not everyone through at once

Rate limiting and request limits

limit_req caps requests per second per client IP at the edge, before they cost Node anything; limit_conn caps simultaneous connections; client_max_body_size caps uploads.

In detail

Define zones in the http block and apply them per location: stricter for login and OTP routes, looser for reads. burst lets short spikes through with nodelay. Return 429, and log rejections so you can tune. Edge limits complement, not replace, per user limits in the app with Redis.

limits.conf
nginx
limit_req_zone $binary_remote_addr zone=api:10m rate=20r/s;location /api/ { limit_req zone=api burst=40 nodelay; }
Why it matters A 10 MB zone tracks about 160,000 client IPs, so the memory cost of edge rate limiting is tiny.

Module 20 of 62, phase 03

Telling browsers to be careful

Security headers and hiding details

Add headers that tell browsers to always use HTTPS, refuse framing, sniffing and unexpected scripts, and hide the Nginx version.

In detail

HSTS forces HTTPS for a year once a browser has seen it; only enable it when every subdomain works over HTTPS. X-Content-Type-Options stops MIME sniffing, frame-ancestors in a Content Security Policy blocks clickjacking, and Referrer-Policy limits leaked URLs. A full CSP for the frontend is the strongest XSS defence but needs testing in report only mode first. Block access to dotfiles such as .env and .git.

headers.conf
nginx
add_header Strict-Transport-Security "max-age=31536000; includeSubDomains" always;add_header X-Content-Type-Options nosniff always;
Why it matters The dotfile rule closes the classic leak where a stray .env or .git folder ends up downloadable.

Module 21 of 62, phase 03

Changing the rules without closing

Testing and gracefully reloading Nginx

nginx -t checks syntax and file paths; systemctl reload nginx starts new workers with the new config while old workers finish their in flight requests. No connection is dropped.

In detail

Restart kills everything and briefly refuses connections; reload does not. If the test fails, the running Nginx keeps the old config, so a typo cannot take the site down as long as you always test first. Keep configs in git, deploy them with your app or with Ansible, and reload in the same script that switches upstreams for blue green deploys.

reload.sh
Shell
sudo nginx -t && sudo systemctl reload nginx
Why it matters Because reload only applies a config that passed the test, the worst a typo can do is fail the deploy, not the site.

Module 22 of 62, phase 03

Every visit, written down

Nginx access and error logs

The access log records every request with status, size and timing; the error log records proxy failures, permission problems and config issues. A JSON log format makes both easy to ship to CloudWatch and query.

In detail

Add request_time and upstream_response_time to the format: the difference shows whether slowness is in Nginx and the network or in Node. Include $request_id and pass it to Node so one id links Nginx, app and database logs. Logrotate keeps the files from filling the disk; Ubuntu's package already rotates daily.

logs.conf
nginx
log_format json escape=json '{"time":"$time_iso8601","status":$status,"rt":$request_time}';access_log /var/log/nginx/access.json json;
Why it matters Comparing rt with urt tells you instantly whether a slow request was slow in Node or slow on the client's network.

Phase 04, modules 23 to 30

Keep it running

PM2 process management

Start, cluster, reload with zero downtime, rotate logs and survive reboots.

Module 23 of 62, phase 04

Somebody has to restart it

Why a process manager

node server.js dies when it crashes, when you close SSH, and when the server reboots. A process manager restarts it, runs one copy per CPU core, collects logs and reloads code without downtime.

In detail

PM2 is the common choice for Node on a VPS: it daemonises your app, restarts on crash with backoff, offers cluster mode, graceful reloads, log files with rotation and a boot script. systemd can do the restart and boot part for any program and is a fine alternative for single process apps; containers hand the job to Docker or ECS. On Lambda there is no long running process at all, so none of this applies.

Ways to keep a Node process alive
ToolRestart on crashMulti coreZero downtime reloadLogs
node server.jsNoNoNoTerminal only
nohup or screenNoNoNoOne file, unrotated
systemd unitYesOne process, or many unitsNojournald
PM2Yes, with backoffCluster modepm2 reloadFiles with logrotate
Docker or ECSRestart policyMore containersRolling updatesDriver to CloudWatch

Module 24 of 62, phase 04

The ten commands you need

PM2 essentials

pm2 start launches and names a process; status, logs, restart, reload, stop and delete manage it; describe shows everything about one app.

In detail

Name every process so scripts can target it. pm2 status shows uptime, restarts and memory; a rising restart count means crashes you should read about in pm2 logs. pm2 describe shows the exact script path, node version and environment PM2 is using, which settles most 'works on my machine' puzzles. Run PM2 as the deploy user, never root.

pm2.sh
Shell
pm2 start dist/main.js --name api && pm2 status && pm2 logs api
Why it matters Watch the restart column in pm2 status: a number that keeps climbing is a crash loop hiding behind auto restarts.

Module 25 of 62, phase 04

Config you can commit

The PM2 ecosystem file

ecosystem.config.js declares every process: script, working directory, instances, environment, log paths, memory limits and graceful shutdown timings. Commit it to the repo.

In detail

Point cwd at the current symlink so reloads pick up new releases. Use env_production for production values and load secrets from a file outside the repo with dotenv or from SSM at startup. max_memory_restart restarts a worker that leaks past a threshold; listen_timeout and kill_timeout control how long PM2 waits for a worker to become ready and to shut down cleanly.

ecosystem.config.js
JavaScript
module.exports = { apps: [{ name: "api", script: "dist/main.js", instances: "max", exec_mode: "cluster" }] };
Why it matters wait_ready with listen_timeout is what makes reloads truly zero downtime: PM2 only kills an old worker after the new one says it is ready.

Module 26 of 62, phase 04

One worker per core

Cluster mode

Node runs your JavaScript on one CPU core. PM2 cluster mode starts one worker per core and shares the port between them, so a four core server handles roughly four times the load.

In detail

Workers share nothing in memory, so sessions, caches and rate limits must live in Redis or the database, and in memory timers run once per worker (move cron jobs to a single worker or a queue). Leave a core free if the server also runs Redis or heavy Nginx work. Cluster mode is also what enables pm2 reload, because it can replace workers one at a time.

Back of the envelope

  • Workersone per coreminus one if Redis is local
  • Throughputabout cores times onefor CPU bound code
  • Memoryeach worker separatewatch total RSS
  • Cronworker 0 onlyvia NODE_APP_INSTANCE
cluster.sh
Shell
pm2 start dist/main.js -i max --name api && pm2 scale api 3
Why it matters NODE_APP_INSTANCE is set by PM2 per worker, which is the simplest way to stop a cron job running four times.

Module 27 of 62, phase 04

Swapping engines mid flight

Zero downtime reloads and graceful shutdown

pm2 reload replaces cluster workers one by one: start a new worker, wait until it is ready, tell an old worker to stop, let it finish in flight requests, repeat. Users never see a refused connection.

In detail

Your app must cooperate. On start, call process.send('ready') once the server is listening and the database is connected. On SIGINT, stop accepting new connections, finish current requests, close the database pool and exit within kill_timeout. Without this, reload falls back to killing workers mid request. Restart, unlike reload, stops everything at once and causes a brief outage.

Start, wait for ready, drain the old oneEdgeExternalService
PM2EdgeOld workerExternalNew workerServiceNginxEdge1 start2 ready3 SIGINT4 traffic
main.js
JavaScript
server.listen(3000, () => process.send?.("ready"));process.on("SIGINT", () => server.close(() => db.end().then(() => process.exit(0))));
Why it matters The hard exit timer must be shorter than PM2's kill_timeout, or PM2 kills the worker before it can log a clean exit.

Module 28 of 62, phase 04

Logs that do not fill the disk

PM2 logs and rotation

PM2 writes each app's stdout and stderr to files. pm2-logrotate rotates them by size or date, compresses old ones and keeps a fixed number, so a chatty app never fills the disk.

In detail

Log JSON lines to stdout from the app (pino or winston), let PM2 write them to files with timestamps, rotate with pm2-logrotate, and ship the files to CloudWatch with the agent. pm2 flush empties logs, useful after an incident is resolved. Never log secrets, tokens or full card numbers.

logrotate.sh
Shell
pm2 install pm2-logrotate && pm2 set pm2-logrotate:max_size 50M && pm2 set pm2-logrotate:retain 14
Why it matters Disk full is one of the most common causes of a mysterious outage; rotation with a retain limit removes it entirely.

Module 29 of 62, phase 04

Back up after a reboot

Starting PM2 on boot

pm2 startup generates a systemd unit that starts PM2 at boot for your user, and pm2 save records the current process list so it is restored automatically.

In detail

Run pm2 startup as the deploy user, copy the sudo command it prints, then start your apps and run pm2 save. Repeat pm2 save after adding or removing apps. Test it with a real reboot during a quiet window; kernel updates will reboot the box eventually anyway.

startup.sh
Shell
pm2 startup systemd -u deploy --hp /home/deploy && pm2 save
Why it matters Forgetting pm2 save is the classic reason an app is missing after the first reboot.

Module 30 of 62, phase 04

A dashboard in the terminal

Monitoring processes and memory

pm2 monit shows live CPU, memory and logs per worker; pm2 status shows restarts; max_memory_restart recycles leaking workers; and Node's own metrics tell you about the event loop.

In detail

Watch three numbers: restart count (crashes), memory per worker over days (leaks) and event loop lag (blocked CPU). Expose a /metrics or /health endpoint with process.memoryUsage and loop delay, and scrape it from CloudWatch or an uptime checker. A memory restart is a safety net, not a fix; take a heap snapshot to find the leak.

metrics.js
JavaScript
const h = monitorEventLoopDelay(); h.enable();app.get("/metrics", (_, res) => res.json({ rss: process.memoryUsage().rss, lagMs: h.mean / 1e6 }));
Why it matters Event loop lag is the single best signal that synchronous code, not traffic, is slowing a Node API.

Phase 05, modules 31 to 36

Releases without fear

Zero downtime deploys, CI/CD and rollbacks

Atomic switches, blue green on one box, safe migrations, automated pipelines and a fast way back.

Module 31 of 62, phase 05

One switch, no half states

Atomic deploy script with releases

A deploy script builds or uploads into a new release folder, installs dependencies, runs checks, flips the current symlink, reloads PM2 and verifies health, rolling back automatically if the check fails.

In detail

Everything slow or risky happens before the switch; the switch itself is one system call. Shared files are linked into the release. After reload, poll the health endpoint for a few seconds; on failure, point current back to the previous release and reload again. Keep the last five releases and prune the rest.

release.sh
Shell
ln -sfn "$NEW" current.tmp && mv -Tf current.tmp current && pm2 reload apicurl -fsS localhost:3000/health || rollback
Why it matters The script decides on rollback by itself, so a bad deploy at 2 a.m. heals within twenty seconds without a human.

Module 32 of 62, phase 05

Two stages, one spotlight

Blue green on a single server

Run two copies of the app on different ports (blue on 3001, green on 3002). Deploy to the idle one, test it on its port, then point the Nginx upstream at it and reload. The old colour stays warm for instant rollback.

In detail

This works even for apps that cannot use PM2 cluster reloads, such as a single process with long startup, and lets you smoke test the new version on the real server before any user sees it. It costs double memory during the switch. Across several servers the same idea becomes an AWS target group swap or a DNS weight change.

Deploy to the idle colour, then flip NginxEdgeServiceClientData
Nginx upstreamEdgeBlue :3001 (live)ServiceGreen :3002 (new)ClientShared RDSDatanowafter switch
switch.sh
Shell
echo "server 127.0.0.1:$GREEN_PORT;" > /etc/nginx/conf.d/api_active.inc && nginx -t && systemctl reload nginx
Why it matters Testing the idle colour on its own port before switching means users only ever meet a version that already passed its checks.

Module 33 of 62, phase 05

Changing the schema under a running app

Safe database migrations

During a deploy, old and new code run at the same time for a while, so every migration must work with both. Use expand and contract: add first, migrate data, switch code, remove later.

In detail

Never rename or drop a column in the same deploy that stops using it. Add nullable columns, backfill in batches, then add constraints. Create indexes concurrently in PostgreSQL to avoid locking writes. Run migrations as a separate step before switching traffic, with a lock timeout so a long lock fails fast instead of freezing production. Test migrations against a copy of production data.

migration.sql
SQL
SET lock_timeout = '3s';ALTER TABLE users ADD COLUMN display_name text;          -- expandCREATE INDEX CONCURRENTLY users_display_name_idx ON users (display_name);
Why it matters Spreading one change across three releases means any single deploy can be rolled back without touching the database.

Module 34 of 62, phase 05

Push to deploy, safely

CI/CD with GitHub Actions

On every push, CI installs, lints, tests and builds; on main, it packages an artifact and deploys over SSH to the server or uploads to Lambda and S3, then runs a smoke test.

In detail

Build once in CI and ship the artifact; never build on production servers. Store the SSH key and AWS credentials as repository secrets, or better, use GitHub's OIDC to assume an AWS role without long lived keys. Protect main with required checks, use environments for manual approval to production, and limit concurrency so two deploys never overlap.

Build once, deploy the same artifact everywhereClientServiceDataExternal
git pushClientLint, test, buildServiceArtifactDataLinode via SSHServiceLambda + S3Externalon mainrelease.shOIDC role
deploy.yml
YAML
- run: npm ci && npm test && npm run build- run: scp api.tgz deploy@$HOST:/tmp/ && ssh deploy@$HOST "/srv/api/release.sh /tmp/api.tgz"
Why it matters concurrency: production queues a second push behind the first, so two releases never race to flip the same symlink.

Module 35 of 62, phase 05

The fastest fix is going back

Rollback strategy

Every deploy needs a way back that takes under a minute: repoint the symlink and reload, flip the blue green upstream, publish the previous Lambda version, or redeploy the previous frontend build.

In detail

Decide rollback triggers in advance: error rate above baseline, health check failures, or a key business metric dropping. Roll back first, investigate second. Database changes are the hard part, which is why expand and contract migrations matter: if the schema is compatible with the previous code, rollback is just code. Practise it, so the commands are muscle memory.

rollback.sh
Shell
ln -sfn "$(ls -1dt /srv/api/releases/* | sed -n 2p)" /srv/api/current && pm2 reload api
Why it matters Having one rollback command per platform, written before the incident, is what keeps a bad deploy to a two minute blip.

Module 36 of 62, phase 05

Are you alive, and are you ready

Health and readiness endpoints

A liveness endpoint answers 'is the process running'; a readiness endpoint answers 'can it serve traffic', checking the database and other critical dependencies. Deploy scripts, load balancers and uptime monitors all poll them.

In detail

Keep /health cheap and dependency free so a slow database does not cause restarts. Make /ready check the database with a short timeout and return 503 while starting up or shutting down. Exclude health routes from auth, rate limits and access logs. On AWS, target groups and Route 53 health checks use the same endpoints.

health.js
JavaScript
app.get("/health", (_, res) => res.send("ok"));app.get("/ready", async (_, res) => (await db.ping(1000)) ? res.send("ready") : res.status(503).end());
Why it matters Returning the git SHA from /ready lets the deploy script confirm the new version is actually the one answering.

Phase 06, modules 37 to 41

Shipping the frontend

Builds, static hosting, CDN and Next.js

Build artifacts, S3 and CloudFront, cache busting, build time config and running Next.js in production.

Module 37 of 62, phase 06

Turning source into files

Frontend build artifacts

A React, Vue or Svelte app builds into a folder of static files: one index.html, hashed JavaScript and CSS bundles and assets. That folder is the whole frontend deployment.

In detail

Build in CI with the same Node version every time, with production mode on so code is minified and dead code removed. The output is pure static files: serve them from Nginx on your server, or from S3 behind CloudFront. Check bundle size in CI; a sudden jump usually means an accidental heavy import. Source maps help debugging but should not be public, or upload them only to your error tracker.

build.sh
Shell
npm ci && npm run build && ls dist/assets
Why it matters Because every bundle name contains a content hash, an old browser tab and a new deploy never fight over the same file.

Module 38 of 62, phase 06

Your frontend on every continent

S3 and CloudFront static hosting

Upload the build to a private S3 bucket and put CloudFront in front: it caches files at edge locations worldwide, serves HTTPS with a free ACM certificate and routes /api to your backend if you want one domain.

In detail

Keep the bucket private and grant CloudFront access with origin access control. Upload hashed assets with a one year cache and index.html with no cache, then invalidate only /index.html after each deploy. For SPA routes, return index.html for 403 and 404 from S3. A second behaviour can forward /api/* to your Linode or an API Gateway, avoiding CORS.

One domain, edge cached files, API behind the same doorClientEdgeDataServiceExternal
BrowserClientCloudFront edgeEdgeS3 bucket (private)DataAPI on Linode or LambdaServiceACM certificateExternalHTTPS/*/api/*
publish-web.sh
Shell
aws s3 sync dist/ s3://acme-web --delete && aws cloudfront create-invalidation --distribution-id E123 --paths "/index.html"
Why it matters Invalidating one path instead of /* is faster and avoids paying for thousands of invalidations each month.

Module 39 of 62, phase 06

New names for new files

Cache busting and cache headers

Give every built file a content hash in its name and cache it forever; keep the HTML that references them uncached. Users always get the latest app, and repeat visits load almost nothing.

In detail

Cache-Control: public, max-age=31536000, immutable for /assets, and no-cache for index.html, which still allows a fast 304 Not Modified check. Service workers add another cache layer that can pin old versions if misconfigured. For API responses, send no-store on personal data and short public max-age only on shared, cacheable data.

Cache headers for each kind of response
ResponseCache-ControlCDN behaviourWhy
/assets/app.3f9c1a.jspublic, max-age=31536000, immutableCached a yearName changes when content changes
/index.htmlno-cacheRevalidated each timeMust always point at the newest bundles
/api/catalog (public)public, max-age=30Short shared cacheAbsorbs spikes, stays fresh
/api/me (personal)private, no-storeNever cachedPersonal data must not be shared
/config.jsonno-storeNever cachedRuntime config can change any time

Module 40 of 62, phase 06

Everything in the bundle is public

Frontend environment variables

Frontend environment variables are inlined into JavaScript at build time and anyone can read them. Only public values belong there: API base URLs, analytics keys, feature flags.

In detail

Vite exposes only variables prefixed with VITE_, Next.js only NEXT_PUBLIC_. Build once per environment, or, to build once and deploy everywhere, serve a small /config.json from the server that the app fetches at startup. Secrets such as database URLs or private API keys must stay on the server, behind your own API.

runtime-config.js
JavaScript
const config = await fetch("/config.json").then((r) => r.json());
Why it matters Runtime config lets the exact same tested build move from staging to production, which removes a whole class of build drift bugs.

Module 41 of 62, phase 06

When the frontend has a server too

Deploying Next.js with PM2 and Nginx

Next.js with server rendering is a Node app, not static files. Build with output: 'standalone', run the generated server with PM2, put Nginx in front for TLS and caching of /_next/static.

In detail

The standalone output contains a minimal server.js and only the node_modules it needs, so releases are small. Copy .next/static and public alongside it. Nginx serves /_next/static with a long immutable cache and proxies everything else. A fully static export (output: 'export') can go to S3 and CloudFront instead.

next.sh
JavaScript
next build && cp -r .next/static .next/standalone/.next/ && pm2 start .next/standalone/server.js --name web
Why it matters Letting Nginx serve /_next/static directly keeps Node workers free for the pages that actually need rendering.

Phase 07, modules 42 to 52

The AWS you actually use

IAM, EC2, RDS, Lambda, S3 and Secrets

A short, practical tour of the AWS services a typical Node product really touches, set up safely.

Module 42 of 62, phase 07

Keys that open only one door

AWS accounts, IAM and least privilege

IAM decides who can do what in AWS. Lock the root account away with MFA, give people their own users or SSO logins, and give apps roles with only the permissions they need.

In detail

Never use root for daily work or put access keys in code. EC2 instances and Lambda functions get IAM roles, so credentials rotate automatically. CI uses OIDC to assume a deploy role. Write policies that name exact actions and resources, for example s3:PutObject on one bucket prefix, instead of s3:*. Turn on CloudTrail so every API call is recorded, and set a budget alarm on day one.

policy.json
JSON
{ "Effect": "Allow", "Action": ["s3:PutObject"], "Resource": "arn:aws:s3:::acme-uploads/avatars/*" }
Why it matters Scoping each statement to one bucket prefix, one secret path and one log group limits the damage if the app is ever compromised.

Module 43 of 62, phase 07

Same server, different landlord

EC2 and Linode side by side

An EC2 instance and a Linode are both virtual servers you SSH into and set up the same way. The difference is everything around them: VPCs, security groups, IAM roles and managed services on AWS versus simplicity and predictable pricing on Linode.

In detail

Choose Linode when you want one or a few servers, fixed monthly bills and a simple dashboard. Choose EC2 when the app leans on other AWS services (RDS, S3, Lambda, SQS) and should reach them over a private network with IAM roles instead of keys. Mixing is fine: a Linode API can use RDS over TLS, though cross provider traffic adds latency and egress cost.

EC2 and Linode compared
AspectAWS EC2Linode
PricingPer second, many line itemsFlat monthly, bandwidth included
NetworkingVPC, subnets, security groups, Elastic IPPublic and private IPs, Cloud Firewall, VLANs
CredentialsIAM instance roles, no keysAPI tokens, IAM user keys for AWS
Managed services nearbyRDS, S3, Lambda, SQS in the same VPCManaged databases, Object Storage
Load balancerApplication Load BalancerNodeBalancer
Learning curveSteeperGentle

Module 44 of 62, phase 07

A database someone else patches

Amazon RDS for PostgreSQL

RDS runs PostgreSQL for you: installation, patching, automated backups, point in time recovery, monitoring and optional Multi AZ failover. You pick the instance size, storage and network placement.

In detail

Start with a burstable db.t4g instance for small apps, gp3 storage with autoscaling, a parameter group for settings such as statement_timeout and slow query logging, and Performance Insights on. Multi AZ keeps a synchronous standby in another zone and fails over in about a minute; read replicas scale reads. Size connections: each Node worker holds a pool, so total connections equal servers times workers times pool size.

Back of the envelope

  • Starter instancedb.t4g.small2 vCPU burstable, 2 GB
  • Max connectionsabout 200at 2 GB, depends on memory
  • Backups7 to 35 days PITRautomated
  • Multi AZ failoverabout 60 to 120 sDNS flips to standby
pool.js
JavaScript
const pool = new Pool({ connectionString: process.env.DATABASE_URL, max: 10, ssl: { rejectUnauthorized: true } });
Why it matters Verifying the RDS certificate with the AWS CA bundle stops a man in the middle, which plain ssl: true does not.

Module 45 of 62, phase 07

No public door to the data

RDS networking, security groups and TLS

Put RDS in private subnets with public access off, allow port 5432 only from the app's security group, and require TLS. Connect from your laptop through a bastion, SSM Session Manager or a VPN, never by opening the database to the internet.

In detail

Security groups can reference other security groups, so 'allow from sg-app' keeps working as servers come and go. If the app runs on Linode outside AWS, allow only the Linode's static IP and enforce rds.force_ssl. Rotate the master password through Secrets Manager and create a separate, least privileged database user for the app.

Only the app's security group and an audited tunnel reach the databaseClientEdgeServiceData
Your laptopClientSSM tunnelEdgeApp servers (sg-app)ServiceRDS private subnetData5432 TLSport forward
tunnel.sh
Shell
aws ssm start-session --target i-0abc --document-name AWS-StartPortForwardingSessionToRemoteHost \  --parameters host=db.xxxx.rds.amazonaws.com,portNumber=5432,localPortNumber=5433
Why it matters Session Manager tunnels need no open SSH port and leave an audit trail, which beats a public bastion host.

Module 46 of 62, phase 07

Rewind to any second

RDS backups, snapshots and restores

RDS takes daily snapshots and keeps transaction logs, so you can restore to any second within the retention window. Manual snapshots stay until you delete them; take one before risky migrations.

In detail

A restore always creates a new instance; you then point the app at it or rename. Practise the restore and time it. Copy snapshots to another region for disaster recovery, and enable deletion protection on production instances. For long term or cross provider safety, add a nightly pg_dump to S3.

restore.sh
Shell
aws rds restore-db-instance-to-point-in-time --source-db-instance-identifier prod \  --target-db-instance-identifier prod-restore --restore-time 2026-10-06T08:31:00Z
Why it matters Restores create a new instance, so copy the security group and subnet group explicitly or the app will not be able to reach it.

Module 47 of 62, phase 07

Code that only runs when called

AWS Lambda for Node

Lambda runs a function in response to an event: an HTTP request through API Gateway, a file landing in S3, a message on SQS, or a schedule. You pay per request and per millisecond, and it scales to zero.

In detail

Lambda suits spiky or background work: image resizing, webhooks, scheduled reports, queue consumers. Keep handlers small, create clients and database connections outside the handler so warm invocations reuse them, and set memory (which also sets CPU) by measuring. Cold starts add a few hundred milliseconds for Node; provisioned concurrency removes them for latency critical paths. The 15 minute limit rules out long jobs.

Back of the envelope

  • Timeout limit15 minutesper invocation
  • Memory128 MB to 10 GBCPU scales with it
  • Node cold startabout 200 to 600 mssmaller bundles start faster
  • Free tier1 M requests per monthplus 400k GB seconds
handler.mjs
JavaScript
export const handler = async (event) => ({ statusCode: 200, body: JSON.stringify({ ok: true }) });
Why it matters Writing thumbnails to a different prefix than the trigger avoids the classic infinite loop of a Lambda triggering itself.

Module 48 of 62, phase 07

Versions you can point at

Deploying Lambda with API Gateway, versions and aliases

Bundle the function with esbuild, upload it, publish an immutable version, and move a live alias to it. API Gateway invokes the alias, so rollback is moving the alias back.

In detail

Use the AWS SAM or Serverless Framework, or plain CLI commands in CI. HTTP APIs in API Gateway are cheaper and simpler than REST APIs for most cases. Weighted aliases allow canary releases: send 10 percent to the new version, watch CloudWatch errors, then shift the rest. Set reserved concurrency to protect your database from a sudden burst.

Shift traffic between immutable versionsClientEdgeService
ClientsClientAPI GatewayEdgeAlias: liveServiceVersion 41 (90%)ServiceVersion 42 (10%)ClientHTTPSinvokecanary
deploy-lambda.sh
Shell
aws lambda update-function-code --function-name orders-api --zip-file fileb://fn.zip --publishaws lambda update-alias --function-name orders-api --name live --function-version $V
Why it matters Because API Gateway calls the live alias, promotion and rollback are both a single alias update.

Module 49 of 62, phase 07

Many functions, few connections

Lambda with RDS through RDS Proxy

Each Lambda instance opens its own database connection, so a traffic spike of 1,000 concurrent functions can exhaust PostgreSQL. RDS Proxy pools and shares connections, survives failovers and can use IAM authentication.

In detail

Put the Lambda in the same VPC private subnets as the proxy, give it a security group allowed by the proxy, and connect to the proxy endpoint instead of the database. Keep the pg client outside the handler with a pool size of 1. Set reserved concurrency so even the proxy is not overwhelmed. For light, bursty workloads, the RDS Data API is an alternative over HTTPS.

db-lambda.mjs
JavaScript
const pool = new pg.Pool({ host: process.env.PROXY_HOST, max: 1 });export const handler = async () => (await pool.query("select now()")).rows[0];
Why it matters IAM tokens through the proxy mean no database password lives in Lambda configuration at all.

Module 50 of 62, phase 07

Files that never fill your disk

S3 for user uploads and assets

Store user files in S3, not on the server's disk, so any server or Lambda can reach them and deploys never lose them. Let browsers upload directly with pre signed URLs.

In detail

The API creates a short lived pre signed PUT URL for a specific key and content type; the browser uploads straight to S3; an S3 event triggers a Lambda for thumbnails or virus scanning. Keep buckets private with Block Public Access on, serve files through CloudFront or pre signed GET URLs, turn on versioning, and add lifecycle rules to move old files to cheaper storage.

presign.js
JavaScript
const url = await getSignedUrl(s3, new PutObjectCommand({ Bucket, Key, ContentType }), { expiresIn: 300 });
Why it matters Uploads bypass your server completely, so a 50 MB video costs your Node process nothing.

Module 51 of 62, phase 07

Passwords out of the code

Secrets Manager and Parameter Store

Keep database passwords, API keys and signing secrets in AWS Secrets Manager or SSM Parameter Store, read them at startup with the instance or Lambda role, and never commit them or bake them into images.

In detail

Parameter Store SecureString is free for standard parameters and fine for most config; Secrets Manager costs a little but adds automatic rotation, notably for RDS passwords. On a Linode outside AWS, a root owned .env with permissions 600 in the shared folder is acceptable, or fetch from AWS with a tightly scoped access key. Rotate anything that was ever pasted into chat or a ticket.

secrets.js
JavaScript
const { SecretString } = await sm.send(new GetSecretValueCommand({ SecretId: "prod/api/db" }));
Why it matters Loading secrets into memory at startup through the instance role means nothing sensitive sits in files, git or PM2 config.

Module 52 of 62, phase 07

No surprise invoices

Budgets, alarms and cost hygiene

Set an AWS Budget with email alerts the day you create the account, tag resources by project, and review Cost Explorer monthly. Most surprise bills come from forgotten instances, NAT gateways, data transfer and unlimited log retention.

In detail

Right size instances after watching real usage, use Graviton (t4g, m7g) for cheaper compute, set CloudWatch log retention instead of forever, move old S3 objects to infrequent access, and delete unattached volumes and old snapshots. A NAT gateway costs money every hour plus per gigabyte; VPC endpoints for S3 avoid some of it. Linode's flat pricing is easier to predict for steady workloads.

budget.sh
Shell
aws budgets create-budget --account-id 123456789012 --budget file://budget.json --notifications-with-subscribers file://notify.json
Why it matters A forecasted alert warns you mid month, while there is still time to fix the cause, not after the bill arrives.

Phase 08, modules 53 to 59

Seeing what happens

Logs, CloudWatch, alerts and debugging

Structured logs, shipping them to CloudWatch, queries, alarms, uptime checks and a calm incident routine.

Module 53 of 62, phase 08

Logs a machine can read

Structured JSON logging in Node

Log one JSON object per line with a level, message, request id and useful fields. JSON logs can be searched, filtered and graphed in CloudWatch Logs Insights; plain sentences cannot.

In detail

Use pino for speed. Attach a request id (from Nginx's X-Request-Id) to every line via a child logger, log the start and end of each request with status and duration, and log errors with their stack. Redact authorization headers, passwords and tokens at the logger level. Use levels consistently: info for business events, warn for recoverable problems, error for failures that need a look.

logger.js
JavaScript
const logger = pino({ level: "info", redact: ["req.headers.authorization"] });
Why it matters Reusing Nginx's request id means one search finds the same request in the proxy log, the app log and the slow query log.

Module 54 of 62, phase 08

Logs that outlive the server

Shipping logs and metrics with the CloudWatch agent

The CloudWatch agent tails Nginx and PM2 log files and sends them to CloudWatch Logs, and also reports memory and disk metrics that EC2 does not collect by default. It works on Linode too, with an IAM user key.

In detail

Create one log group per source (/app/api, /nginx/access, /nginx/error), set retention, and use the instance id or hostname as the stream name. On EC2 attach the CloudWatchAgentServerPolicy to the instance role; on Linode use a dedicated IAM user limited to those log groups. Lambda sends its console output to CloudWatch automatically.

Servers and functions land in one placeEdgeServiceExternal
Nginx JSON logsEdgePM2 app logsServiceCloudWatch agentEdgeCloudWatch LogsExternalLambda console.logServicetailtailshipautomatic
amazon-cloudwatch-agent.json
JSON
{ "logs": { "logs_collected": { "files": { "collect_list": [{ "file_path": "/srv/api/shared/logs/api.out.log", "log_group_name": "/app/api" }] } } } }
Why it matters Memory and disk usage are not EC2 default metrics; the agent is how you get an alarm before the disk fills up.

Module 55 of 62, phase 08

Asking your logs questions

CloudWatch Logs Insights queries

Logs Insights queries JSON logs across groups with a small pipe language: filter, stats, sort. Find the slowest routes, count 5xx by path, or follow one request id across Nginx and the app.

In detail

Save the handful of queries you reach for during incidents. Because logs are JSON, fields such as status, rt and route are queryable without parsing. Queries are billed by data scanned, so narrow the time range first. Metric filters can turn a log pattern (such as level error) into a CloudWatch metric for alarms.

queries.txt
Notes
fields @timestamp, uri, status, rt | filter status >= 500 | stats count() by uri | sort count() desc
Why it matters Querying p99 rather than average shows the slow tail that real users complain about.

Module 56 of 62, phase 08

Being told before users tell you

CloudWatch metrics, alarms and SNS

Alarms watch a metric and notify you through SNS by email, SMS or a chat webhook when it crosses a threshold for a few minutes: high 5xx rate, RDS CPU or free storage, Lambda errors, disk usage, or a missing heartbeat.

In detail

Alert on symptoms users feel (errors, latency, availability) more than on causes (CPU). Use several evaluation periods to avoid flapping and treat missing data as breaching for heartbeats. Start with a short list and tune it: an alarm that fires weekly for nothing trains people to ignore it.

A starter set of alarms
AlarmMetricThresholdWhy it matters
API errors spikeAppErrors from a log metric filterover 20 per minute, 3 of 5 minutesUsers are seeing failures
API not readyRoute 53 health check on /ready3 failed checksSite down from outside
RDS CPU highAWS/RDS CPUUtilizationover 80% for 15 minutesSlow queries ahead
RDS storage lowAWS/RDS FreeStorageSpaceunder 5 GBWrites will fail when it hits zero
Disk almost fullApp/Host disk used_percentover 80%Logs or uploads filling the server
Lambda errorsAWS/Lambda Errorsover 1% of invocationsBackground work failing silently
Budget forecastAWS Budgetsover 100% forecastBill surprise ahead
alarms.sh
Shell
aws cloudwatch put-metric-alarm --alarm-name rds-cpu-high --metric-name CPUUtilization --namespace AWS/RDS \  --threshold 80 --comparison-operator GreaterThanThreshold --evaluation-periods 3 --period 300 --statistic Average
Why it matters Three of five datapoints breaching filters out one noisy minute while still catching a real problem within five.

Module 57 of 62, phase 08

A friend who checks from outside

Uptime checks and status from the user's side

An external uptime check requests your site from several locations every minute and alerts if it fails or gets slow. It catches what internal monitoring misses: DNS mistakes, expired certificates, firewall changes, a dead server.

In detail

Check the homepage, the API /ready endpoint and one critical flow. Route 53 health checks, UptimeRobot, Better Stack or Checkly all work. Monitor certificate expiry too. Point alerts at a channel humans actually watch, and give each alert a link to the runbook.

uptime.sh
Shell
aws route53 create-health-check --caller-reference $(date +%s) \  --health-check-config Type=HTTPS,FullyQualifiedDomainName=api.example.com,ResourcePath=/ready,RequestInterval=30
Why it matters Checking from outside every 30 seconds in several regions catches the outages that every internal metric would show as perfectly normal.

Module 58 of 62, phase 08

Calm detective work

Debugging a live server

When production misbehaves, check in order: is the process up (pm2 status), is Nginx happy (error log), are requests failing (access log status counts), is the box healthy (CPU, memory, disk), and is the database slow or out of connections.

In detail

Use read only commands first and change one thing at a time. ss -tlnp shows what listens on which port, htop shows CPU and memory per process, df -h and du find a full disk, journalctl shows system logs, and pg_stat_activity shows long running or blocked queries. Write down what you saw and did with timestamps; it becomes the incident timeline.

triage.sh
Shell
pm2 status; sudo tail -n 50 /var/log/nginx/error.log; df -h; free -m; uptime
Why it matters Running the same short list every time stops panic from skipping the obvious cause, which is usually a full disk or a crashed process.

Module 59 of 62, phase 08

Steps written before you need them

Incident runbook and postmortems

A runbook is a short page per alert: what it means, how to confirm it, the first safe actions, how to roll back, and who to call. A blameless postmortem afterwards turns each incident into a fix.

In detail

Store runbooks next to the code and link them from every alarm. Write for someone tired at 3 a.m.: exact commands, expected output, and when to escalate. Afterwards record the timeline, impact, root cause, what went well and action items with owners. Most action items are small: a missing alarm, a better health check, a migration rule.

runbook-api-5xx.md
Notes
# API 5xx spike1. Check pm2 status and pm2 logs api --err2. If last deploy < 30 min ago: ./rollback.sh
Why it matters Putting rollback as step one for recent deploys removes the most common incident in under a minute, before any deep debugging.

Phase 09, modules 60 to 62

Ship it like you mean it

Hardening and the full reference deployment

Patching, a production checklist and one complete deployment that uses every module.

Module 60 of 62, phase 09

Patching without drama

OS updates, unattended upgrades and reboots

Turn on unattended security upgrades, apply other updates weekly in a quiet window, and reboot when the kernel requires it. PM2 startup and Nginx's systemd unit bring everything back.

In detail

Ubuntu's unattended-upgrades installs security patches daily; enable automatic reboots at a fixed early morning time, or reboot manually after checking /var/run/reboot-required. With two servers behind a load balancer, update one at a time. Keep Node updated to the latest patch of your LTS line and run npm audit in CI. Snapshots before major upgrades make them reversible.

updates.sh
Shell
sudo apt-get install -y unattended-upgrades && sudo dpkg-reconfigure -plow unattended-upgrades
Why it matters Scheduling automatic reboots for 04:30 local time means kernel patches land within a day without anyone staying up.

Module 61 of 62, phase 09

Nothing forgotten on launch day

The production launch checklist

Before real users arrive, walk through security, reliability, observability and recovery. Each item links to the module that explains it, and none takes more than an hour.

Module 62 of 62, phase 09

Everything, working together

The full reference deployment

One complete, realistic setup: React on S3 and CloudFront, a Node API on two Linodes or EC2 instances behind Nginx with PM2 in cluster mode, RDS PostgreSQL in a private subnet, S3 for uploads, Lambda for image processing, secrets in AWS, logs and alarms in CloudWatch, and deploys from GitHub Actions.

In detail

Walk one request through it: DNS to CloudFront for the app shell, then /api to Nginx, which proxies to a PM2 worker, which queries RDS through a pooled TLS connection and logs JSON with the request id. An avatar upload goes straight to S3 with a pre signed URL and a Lambda makes the thumbnail. A deploy builds once in CI, releases atomically with a health checked symlink switch, and can roll back in one command.

Every module, one diagramClientEdgeServiceDataExternal
UsersClientCloudFront + S3 webEdgeNginx x2EdgePM2 cluster APIServiceRDS PostgreSQLDataS3 uploadsDataLambda resizeServiceCloudWatchExternalapp shell/apiproxyTLS poolpresigneventlogslogs

Back of the envelope

  • Monthly costabout 60 to 120 USD2 small servers, db.t4g.small
  • Deploy timeabout 3 minutesCI build + release
  • Rollbackunder 1 minutesymlink or alias
  • RPOseconds to 5 minutesRDS point in time
README-deploy.md
Notes
Web: S3 + CloudFront | API: Nginx -> PM2 cluster | DB: RDS private | Logs: CloudWatch
Why it matters A one page map like this is the document every new teammate and every 3 a.m. incident responder reads first.

Questions people ask

Short answers to the questions that come up most when deploying a Node app, each linked to the module that covers it in depth.

Do I need Nginx if Node can listen on port 443 itself?

You can, but Nginx handles TLS, compression, static files, slow clients, rate limits and graceful config reloads far better, and it lets PM2 reload Node workers behind it without dropping connections.

Module 13: Reverse proxy to Node
PM2 or systemd?

systemd is enough for a single process app. PM2 adds cluster mode, zero downtime reloads and log rotation for Node, which is why most Node VPS setups use it, often started by systemd at boot.

Module 23: Why a process manager
What is the difference between pm2 restart and pm2 reload?

restart stops all workers and starts them again, a brief outage. reload replaces cluster workers one at a time and waits for each to be ready, so no request is dropped.

Module 27: Zero downtime reloads and graceful shutdown
Why do I get 502 Bad Gateway?

Nginx could not reach Node: the app crashed, is still starting, listens on a different port, or is bound to another interface. Check pm2 status, pm2 logs and the Nginx error log.

Module 58: Debugging a live server
Linode or AWS?

Linode for a few servers with predictable bills and a simple setup. AWS when you rely on RDS, S3, Lambda and IAM roles in one private network. Mixing them is fine.

Module 43: EC2 and Linode side by side
When should I use Lambda instead of a server?

For spiky or background work: webhooks, image processing, scheduled jobs and queue consumers. Keep the main API with long lived connections and sockets on a server.

Module 47: AWS Lambda for Node
How do I stop Lambda from exhausting RDS connections?

Put RDS Proxy in front of the database, keep one connection per function instance outside the handler, and set reserved concurrency.

Module 49: Lambda with RDS through RDS Proxy
Where should secrets live?

In AWS Secrets Manager or SSM Parameter Store read through IAM roles, or in a root owned .env file outside the repo on a plain VPS. Never in git, images or frontend bundles.

Module 51: Secrets Manager and Parameter Store
How do I deploy without downtime?

Build an artifact in CI, unpack it into a new release folder, install dependencies, switch the symlink, run pm2 reload with graceful shutdown, health check, and roll back automatically if it fails.

Module 31: Atomic deploy script with releases
Why do users still see the old frontend after a deploy?

index.html is being cached. Serve it with no-cache, give assets hashed names with long caches, and invalidate only /index.html on CloudFront.

Module 39: Cache busting and cache headers
How much logging is too much?

Log one structured line per request plus business events and errors. Set retention in CloudWatch so you do not pay to keep debug noise forever.

Module 53: Structured JSON logging in Node
How is progress saved on this page?

In this browser only, in local storage. Nothing is sent anywhere.

Module 61: The production launch checklist
What keyboard shortcuts does this page support?

T switches theme, F toggles full screen, M opens the menu, slash searches, J and K move between modules, and Esc closes the menu or leaves focus mode.

Module 00: The deployment dictionary

Official documentation and sources

Every source linked from the modules above, grouped by the phase that uses it and then by where it lives. 174 links in total, all opening in a new tab.

Deployment foundations

Phase 01, From laptop to the internet

14

A Linode VPS from zero

Phase 02, Your first server

18

Nginx as reverse proxy and web server

Phase 03, The front door

33

PM2 process management

Phase 04, Keep it running

21

Zero downtime deploys, CI/CD and rollbacks

Phase 05, Releases without fear

16

Builds, static hosting, CDN and Next.js

Phase 06, Shipping the frontend

12

IAM, EC2, RDS, Lambda, S3 and Secrets

Phase 07, The AWS you actually use

32

Logs, CloudWatch, alerts and debugging

Phase 08, Seeing what happens

20

Tools and libraries1

Credits

The technologies this roadmap teaches and the tools used to build the page. The people behind it are listed in the footer.

Keyboard shortcuts

Modules

←→
Previous, next module in focus mode
JK
Next, previous module
O
Focus on the current module
I
Show or hide the details
C
Collapse or expand the module
D
Mark the module as done

Page

M
Module menu
/
Search modules
T
Light or dark theme
F
Full screen
G
Back to the top

Help

?
Open this list
Esc
Close a dialog, the menu or focus mode

Shortcuts pause while you type in a search box.