Backend cheatsheetNestJS observabilityKnow before the pager does
NestJS
Observed
Your service already knows when it is slow, failing or backed up. This guide wires a NestJS app from nothing to metrics, logs, traces, dashboards and alerts, and shows which signal answers which question.
Three signals
Metrics say something is wrong, traces say where, logs say why. Emit all three and join them with one id.
Picture it: from the app to one screen
Metrics are pulled, logs and traces are pushed. The coral path is how you hear about trouble before users do.
| Signal | Answers | Cost per request | Keep for |
|---|---|---|---|
| Metrics | Is it broken? How often, how slow? | Almost nothing, it is a counter | Months |
| Traces | Where did the time go? | One span per step, often sampled | Days |
| Logs | Why did this one fail? | One line per event | Weeks |
Metrics with prom-client
One registry, the Node.js default metrics, an HTTP histogram and a /metrics route for Prometheus to scrape.
Only goes up. Requests served, jobs failed.
See the exampleGoes up and down. Connected sockets, queue depth.
See the exampleCounts values into buckets so you can read percentiles.
See the examplePercentiles computed in the app. Cannot be summed across pods.
See the example@Injectable()
export class MetricsService {
readonly registry = new Registry();
readonly httpDuration = new Histogram({
name: 'http_server_duration_seconds',
help: 'HTTP request duration in seconds',
labelNames: ['method', 'route', 'status'] as const,
buckets: [0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5],
registers: [this.registry],
});
constructor() {
collectDefaultMetrics({ register: this.registry }); // CPU, memory, GC, event loop lag
}
}
Why it matters: buckets placed around your real latencies make histogram_quantile give an honest p95.
@Injectable()
export class HttpMetricsInterceptor implements NestInterceptor {
constructor(private readonly metrics: MetricsService) {}
intercept(ctx: ExecutionContext, next: CallHandler) {
if (ctx.getType() !== 'http') return next.handle();
const http = ctx.switchToHttp();
const req = http.getRequest();
const route = req.route?.path ?? 'unmatched'; // the template, never the raw URL
const end = this.metrics.httpDuration.startTimer({ method: req.method, route });
return next.handle().pipe(tap({
next: () => end({ status: http.getResponse().statusCode }),
error: (err) => end({ status: err?.getStatus?.() ?? 500 }),
}));
}
}
Why it matters: one global interceptor times every controller with three small labels.
Picture it: one request lands in the histogram
183 msobserve(0.183)→.1.25.51+InfEvery bucket at or above the value counts it. That is why buckets are cumulative.Labels and cardinality
Every unique mix of label values is its own time series. Keep labels to small, fixed sets.
A histogram with 8 buckets writes 11 series per label mix: one per bucket, +Inf, plus sum and count. Multiply by routes, methods, statuses and pods and it grows fast. Add a label like userId and it grows without limit. Run this to see the numbers.
const perHistogram = 8 + 1 + 2; // buckets, +Inf, sum and count
const base = { routes: 40, methods: 3, statuses: 6, pods: 4 };
const count = (labels) => Object.values(labels).reduce((a, b) => a * b, perHistogram);
console.log('route, method, status, pod:', count(base).toLocaleString('en-US'));
console.log('plus a raw URL instead of the route:', count({ ...base, routes: 25000 }).toLocaleString('en-US'));
console.log('plus a userId label:', count({ ...base, users: 50000 }).toLocaleString('en-US'));
$ node series-count.js route, method, status, pod: 31,680 plus a raw URL instead of the route: 19,800,000 plus a userId label: 1,584,000,000
Why it matters: one careless label turns about 32 thousand series into millions or billions, and Prometheus runs out of memory.
| Label | Values | Verdict |
|---|---|---|
route template | tens | Keep |
method, status | a handful | Keep |
raw url | unbounded | Drop |
userId, gameId | unbounded | Put on logs and spans |
Structured logs
JSON lines with a request id on every entry, secrets redacted, and Nest’s own logger routed through the same pipe.
LoggerModule.forRoot({
pinoHttp: {
level: process.env.LOG_LEVEL ?? 'info',
genReqId: (req, res) => {
const id = req.headers['x-request-id'] ?? randomUUID();
res.setHeader('x-request-id', id);
return id;
},
redact: ['req.headers.authorization', 'req.headers.cookie', '*.password'],
customProps: () => ({ service: 'game-api' }),
transport: process.env.NODE_ENV !== 'production' ? { target: 'pino-pretty' } : undefined,
},
}),
Why it matters: one id ties every line of a request together, and tokens never reach the log store.
const app = await NestFactory.create(AppModule, { bufferLogs: true });
app.useLogger(app.get(Logger));
Why it matters: startup messages wait for Pino, so nothing prints in two formats.
Traces with OpenTelemetry
Start the SDK before Nest loads, let auto instrumentation cover HTTP, PostgreSQL and Redis, and name the steps that are yours.
Picture it: one slow request as a waterfall
The widest bar is where the time went. Here it is one query, not the code around it.
const sdk = new NodeSDK({
resource: resourceFromAttributes({ [ATTR_SERVICE_NAME]: 'game-api' }),
traceExporter: new OTLPTraceExporter(), // reads OTEL_EXPORTER_OTLP_ENDPOINT
instrumentations: [getNodeAutoInstrumentations({
'@opentelemetry/instrumentation-fs': { enabled: false },
})],
});
sdk.start(); // must run before express, pg or ioredis are imported
process.on('SIGTERM', () => sdk.shutdown().finally(() => process.exit(0)));
Why it matters: HTTP, pg, ioredis and Nest handlers get spans with no code changes, and Pino lines gain trace_id.
async score(move: MoveDto) {
return tracer.startActiveSpan('score-move', async (span) => {
span.setAttributes({ 'game.id': move.gameId, 'word.length': move.word.length });
try {
return await this.computeScore(move);
} catch (err) {
span.recordException(err);
span.setStatus({ code: SpanStatusCode.ERROR });
throw err;
} finally {
span.end();
}
});
}
Why it matters: the game id rides on the span, where one value costs one attribute instead of a new series.
Queues and sockets
HTTP metrics stop at the controller. Background jobs and WebSocket connections need their own numbers.
new Gauge({
name: 'bullmq_jobs',
help: 'Jobs by queue and state',
labelNames: ['queue', 'state'] as const,
registers: [registry],
async collect() { // runs only when Prometheus scrapes
const counts = await powerUps.getJobCounts('waiting', 'active', 'delayed', 'failed');
for (const [state, n] of Object.entries(counts)) this.set({ queue: powerUps.name, state }, n);
},
});
worker.on('completed', (job) =>
jobDuration.observe({ queue: job.queueName, name: job.name }, (job.finishedOn! - job.processedOn!) / 1000));
Why it matters: a rising waiting count is the earliest sign workers cannot keep up.
handleConnection() { this.metrics.wsClients.inc({ namespace: '/game' }); }
handleDisconnect() { this.metrics.wsClients.dec({ namespace: '/game' }); }
Why it matters: connected clients per pod shows uneven load, and a sudden drop often means a bad deploy.
Dashboards and alerts
Five PromQL queries cover most dashboards. Page on what users feel, and give each alert time to be sure.
| Panel | PromQL |
|---|---|
| Rate | sum by (route) (rate(http_server_duration_seconds_count[5m])) |
| Errors | sum(rate(http_server_duration_seconds_count{status=~"5.."}[5m])) / sum(rate(http_server_duration_seconds_count[5m])) |
| p95 | histogram_quantile(0.95, sum by (le, route) (rate(http_server_duration_seconds_bucket[5m]))) |
| Event loop | max by (pod) (nodejs_eventloop_lag_p99_seconds) |
| Backlog | sum by (queue) (bullmq_jobs{state="waiting"}) |
groups:
- name: game-api
rules:
- alert: HighErrorRate
expr: sum(rate(http_server_duration_seconds_count{status=~"5.."}[5m])) / sum(rate(http_server_duration_seconds_count[5m])) > 0.02
for: 10m
labels:
severity: page
annotations:
summary: "5xx above 2% for 10 minutes"
runbook: "https://wiki.internal/runbooks/game-api-errors"
Why it matters: a ratio means the same at 3am and at peak, and the runbook link is the first step of the fix.
- Condition trueexpr crosses 2%
- Pendingheld for 10m
- Firingsent to Alertmanager
- Routedgrouped, deduplicated
- Pagedwith the runbook
Which one do I need?
Start from the question you need answered, then reach for the signal that answers it cheapest.
| I want to know | Look at | Tool |
|---|---|---|
| Is the API failing more than usual? | Error ratio | Prometheus |
| Are users waiting too long? | p95 by route | Prometheus |
| Which step made this request slow? | The trace waterfall | Tempo |
| Why did this one move fail? | Log lines for its request id | Loki |
| Are workers keeping up? | Waiting jobs and job duration | Prometheus |
| Is one pod overloaded? | Sockets and event loop lag by pod | Prometheus |
| Should someone wake up? | A symptom alert with for and a runbook | Alertmanager |