NestJS Observed 0/7

Backend cheatsheetNestJS observabilityKnow before the pager does

NestJS
Observed

Your service already knows when it is slow, failing or backed up. This guide wires a NestJS app from nothing to metrics, logs, traces, dashboards and alerts, and shows which signal answers which question.

PrometheusOpenTelemetry7 min read
01

Three signals

Metrics say something is wrong, traces say where, logs say why. Emit all three and join them with one id.

FoundationsMust knowRED methodCorrelation

Picture it: from the app to one screen

NestJS appthree signalsMetrics/metrics, pulledPrometheusLogsstdout JSONLokiTracesOTLP, pushedTempoGrafanaone screenAlertmanagerSlack, pageralert rules fire

Metrics are pulled, logs and traces are pushed. The coral path is how you hear about trouble before users do.

SignalAnswersCost per requestKeep for
MetricsIs it broken? How often, how slow?Almost nothing, it is a counterMonths
TracesWhere did the time go?One span per step, often sampledDays
LogsWhy did this one fail?One line per eventWeeks
02

Metrics with prom-client

One registry, the Node.js default metrics, an HTTP histogram and a /metrics route for Prometheus to scrape.

Metricsprom-clientHistogramInterceptor
Counter.inc()

Only goes up. Requests served, jobs failed.

See the example
rate()
Gauge.set() .inc() .dec()

Goes up and down. Connected sockets, queue depth.

See the example
Now
Histogram.observe(s)

Counts values into buckets so you can read percentiles.

See the example
p95
Summary.observe(s)

Percentiles computed in the app. Cannot be summed across pods.

See the example
Avoid
metrics.service.tsTypeScript
@Injectable()
export class MetricsService {
  readonly registry = new Registry();

  readonly httpDuration = new Histogram({
    name: 'http_server_duration_seconds',
    help: 'HTTP request duration in seconds',
    labelNames: ['method', 'route', 'status'] as const,
    buckets: [0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5],
    registers: [this.registry],
  });

  constructor() {
    collectDefaultMetrics({ register: this.registry }); // CPU, memory, GC, event loop lag
  }
}

Why it matters: buckets placed around your real latencies make histogram_quantile give an honest p95.

http-metrics.interceptor.tsTypeScript
@Injectable()
export class HttpMetricsInterceptor implements NestInterceptor {
  constructor(private readonly metrics: MetricsService) {}

  intercept(ctx: ExecutionContext, next: CallHandler) {
    if (ctx.getType() !== 'http') return next.handle();
    const http = ctx.switchToHttp();
    const req = http.getRequest();
    const route = req.route?.path ?? 'unmatched'; // the template, never the raw URL
    const end = this.metrics.httpDuration.startTimer({ method: req.method, route });
    return next.handle().pipe(tap({
      next: () => end({ status: http.getResponse().statusCode }),
      error: (err) => end({ status: err?.getStatus?.() ?? 500 }),
    }));
  }
}

Why it matters: one global interceptor times every controller with three small labels.

Picture it: one request lands in the histogram

183 msobserve(0.183)→.1.25.51+InfEvery bucket at or above the value counts it. That is why buckets are cumulative.
03

Labels and cardinality

Every unique mix of label values is its own time series. Keep labels to small, fixed sets.

MetricsMust knowSeriesLabels

A histogram with 8 buckets writes 11 series per label mix: one per bucket, +Inf, plus sum and count. Multiply by routes, methods, statuses and pods and it grows fast. Add a label like userId and it grows without limit. Run this to see the numbers.

series-count.jsJavaScript
const perHistogram = 8 + 1 + 2; // buckets, +Inf, sum and count
const base = { routes: 40, methods: 3, statuses: 6, pods: 4 };
const count = (labels) => Object.values(labels).reduce((a, b) => a * b, perHistogram);

console.log('route, method, status, pod:', count(base).toLocaleString('en-US'));
console.log('plus a raw URL instead of the route:', count({ ...base, routes: 25000 }).toLocaleString('en-US'));
console.log('plus a userId label:', count({ ...base, users: 50000 }).toLocaleString('en-US'));
TerminalOutput
$ node series-count.js
route, method, status, pod: 31,680
plus a raw URL instead of the route: 19,800,000
plus a userId label: 1,584,000,000

Why it matters: one careless label turns about 32 thousand series into millions or billions, and Prometheus runs out of memory.

LabelValuesVerdict
route templatetensKeep
method, statusa handfulKeep
raw urlunboundedDrop
userId, gameIdunboundedPut on logs and spans
04

Structured logs

JSON lines with a request id on every entry, secrets redacted, and Nest’s own logger routed through the same pipe.

Logsnestjs-pinoRedactionRequest id
one log linelevel30time1759212000123req.id'r-5f2c'trace_id'4bf92f35...'route'/games/:id/moves'userId'u_91'msg'move rejected'
app.module.tsTypeScript
LoggerModule.forRoot({
  pinoHttp: {
    level: process.env.LOG_LEVEL ?? 'info',
    genReqId: (req, res) => {
      const id = req.headers['x-request-id'] ?? randomUUID();
      res.setHeader('x-request-id', id);
      return id;
    },
    redact: ['req.headers.authorization', 'req.headers.cookie', '*.password'],
    customProps: () => ({ service: 'game-api' }),
    transport: process.env.NODE_ENV !== 'production' ? { target: 'pino-pretty' } : undefined,
  },
}),

Why it matters: one id ties every line of a request together, and tokens never reach the log store.

main.tsTypeScript
const app = await NestFactory.create(AppModule, { bufferLogs: true });
app.useLogger(app.get(Logger));

Why it matters: startup messages wait for Pino, so nothing prints in two formats.

05

Traces with OpenTelemetry

Start the SDK before Nest loads, let auto instrumentation cover HTTP, PostgreSQL and Redis, and name the steps that are yours.

TracesNodeSDKOTLPSpans

Picture it: one slow request as a waterfall

POST /games/:id/movesRolesGuardredis GET sessionpg SELECT wordsscore-movebullmq add0 ms206 ms412 ms

The widest bar is where the time went. Here it is one query, not the code around it.

tracing.tsTypeScript
const sdk = new NodeSDK({
  resource: resourceFromAttributes({ [ATTR_SERVICE_NAME]: 'game-api' }),
  traceExporter: new OTLPTraceExporter(), // reads OTEL_EXPORTER_OTLP_ENDPOINT
  instrumentations: [getNodeAutoInstrumentations({
    '@opentelemetry/instrumentation-fs': { enabled: false },
  })],
});
sdk.start(); // must run before express, pg or ioredis are imported
process.on('SIGTERM', () => sdk.shutdown().finally(() => process.exit(0)));

Why it matters: HTTP, pg, ioredis and Nest handlers get spans with no code changes, and Pino lines gain trace_id.

scoring.service.tsTypeScript
async score(move: MoveDto) {
  return tracer.startActiveSpan('score-move', async (span) => {
    span.setAttributes({ 'game.id': move.gameId, 'word.length': move.word.length });
    try {
      return await this.computeScore(move);
    } catch (err) {
      span.recordException(err);
      span.setStatus({ code: SpanStatusCode.ERROR });
      throw err;
    } finally {
      span.end();
    }
  });
}

Why it matters: the game id rides on the span, where one value costs one attribute instead of a new series.

06

Queues and sockets

HTTP metrics stop at the controller. Background jobs and WebSocket connections need their own numbers.

Async workBullMQSocket.IOGauges
queue.metrics.tsTypeScript
new Gauge({
  name: 'bullmq_jobs',
  help: 'Jobs by queue and state',
  labelNames: ['queue', 'state'] as const,
  registers: [registry],
  async collect() { // runs only when Prometheus scrapes
    const counts = await powerUps.getJobCounts('waiting', 'active', 'delayed', 'failed');
    for (const [state, n] of Object.entries(counts)) this.set({ queue: powerUps.name, state }, n);
  },
});

worker.on('completed', (job) =>
  jobDuration.observe({ queue: job.queueName, name: job.name }, (job.finishedOn! - job.processedOn!) / 1000));

Why it matters: a rising waiting count is the earliest sign workers cannot keep up.

game.gateway.tsTypeScript
handleConnection() { this.metrics.wsClients.inc({ namespace: '/game' }); }
handleDisconnect() { this.metrics.wsClients.dec({ namespace: '/game' }); }

Why it matters: connected clients per pod shows uneven load, and a sudden drop often means a bad deploy.

Queuewaitingactivefailedjob duration
Socketsconnectedevents inevents out
Runtimeevent loop lagheap usedGC time
07

Dashboards and alerts

Five PromQL queries cover most dashboards. Page on what users feel, and give each alert time to be sure.

OperationsPromQLGrafanaAlertmanager
PanelPromQL
Ratesum by (route) (rate(http_server_duration_seconds_count[5m]))
Errorssum(rate(http_server_duration_seconds_count{status=~"5.."}[5m])) / sum(rate(http_server_duration_seconds_count[5m]))
p95histogram_quantile(0.95, sum by (le, route) (rate(http_server_duration_seconds_bucket[5m])))
Event loopmax by (pod) (nodejs_eventloop_lag_p99_seconds)
Backlogsum by (queue) (bullmq_jobs{state="waiting"})
game-api.rules.ymlYAML
groups:
  - name: game-api
    rules:
      - alert: HighErrorRate
        expr: sum(rate(http_server_duration_seconds_count{status=~"5.."}[5m])) / sum(rate(http_server_duration_seconds_count[5m])) > 0.02
        for: 10m
        labels:
          severity: page
        annotations:
          summary: "5xx above 2% for 10 minutes"
          runbook: "https://wiki.internal/runbooks/game-api-errors"

Why it matters: a ratio means the same at 3am and at peak, and the runbook link is the first step of the fix.

  1. Condition trueexpr crosses 2%
  2. Pendingheld for 10m
  3. Firingsent to Alertmanager
  4. Routedgrouped, deduplicated
  5. Pagedwith the runbook

Which one do I need?

Start from the question you need answered, then reach for the signal that answers it cheapest.

I want to knowLook atTool
Is the API failing more than usual?Error ratioPrometheus
Are users waiting too long?p95 by routePrometheus
Which step made this request slow?The trace waterfallTempo
Why did this one move fail?Log lines for its request idLoki
Are workers keeping up?Waiting jobs and job durationPrometheus
Is one pod overloaded?Sockets and event loop lag by podPrometheus
Should someone wake up?A symptom alert with for and a runbookAlertmanager