Backend cheatsheetBullMQ queuesSlow work, off the request
BullMQ
Toolbox
Queues, workers and every state a job passes through, then the options, limits and Redis settings that keep background work reliable once real traffic arrives.
Queue, worker, job
A queue is a name in Redis. Producers add jobs to it, workers take them out and run them. Nothing else is needed to start.
Picture it: producers, Redis, workers
const connection = new IORedis(process.env.REDIS_URL!, { maxRetriesPerRequest: null });
export const powerUps = new Queue('power-ups', { connection });
await powerUps.add('magic-word', { gameId: 'g42', userId: 'u_91' });
new Worker('power-ups', async (job) => {
const hint = await words.findHint(job.data.gameId);
return { hint }; // stored as job.returnvalue
}, { connection, concurrency: 10 });
Why it matters: the API answers at once, and slow work runs on workers you can scale on their own.
The life of a job
Every job moves through a small set of states. Knowing them is how you read the dashboard and debug a stuck queue.
Picture it: every state a job can be in
Not drawn: a flow parent waits in waiting-children until its children finish, job.retry() sends a failed job back to waiting, and a job whose worker died goes back to waiting once its lock expires.
| State | Means | Check with |
|---|---|---|
| waiting | Ready, no worker free yet | getWaitingCount() |
| delayed | Waiting for a time: a delay or a backoff | getDelayedCount() |
| prioritized | Ready, ordered by priority | getPrioritizedCount() |
| active | A worker holds its lock | getActiveCount() |
| completed | Returned a value | getCompletedCount() |
| failed | Threw and has no attempts left | getFailed() |
| waiting-children | A flow parent waiting for its children | getWaitingChildrenCount() |
Adding jobs
Options decide when a job runs, how often it retries, whether duplicates are allowed and how long it is kept.
await emails.add('grievance-report', { date: '2026-09-30' }, {
jobId: 'grievance-report:2026-09-30', // same id twice: the second add is ignored
attempts: 5,
backoff: { type: 'exponential', delay: 2_000 },
delay: 60_000, // not before one minute from now
priority: 1, // 1 is the highest; omit for plain FIFO
removeOnComplete: { age: 86_400, count: 1_000 },
removeOnFail: { age: 7 * 86_400 },
});
await emails.addBulk(recipients.map((r) => ({ name: 'send', data: r }))); // one round trip
Why it matters: a stable jobId makes a nightly job safe to trigger twice, and removeOn options stop Redis from filling up.
How long does a failing job wait between tries? Exponential backoff doubles the delay each time. Run it to see the schedule for the job above.
const delay = 2000;
const attempts = 5;
const exponential = (made) => Math.round(Math.pow(2, made - 1) * delay);
let total = 0;
for (let made = 1; made < attempts; made++) {
const wait = exponential(made);
total += wait;
console.log('after try ' + made + ' wait ' + wait / 1000 + 's');
}
console.log('gives up after ' + attempts + ' tries, ' + total / 1000 + 's of waiting');
$ node backoff.js after try 1 wait 2s after try 2 wait 4s after try 3 wait 8s after try 4 wait 16s gives up after 5 tries, 30s of waiting
Why it matters: five attempts at a 2 second base spread retries over half a minute, enough to ride out a short outage.
| Option | Default | Set it when |
|---|---|---|
attempts | 1 | The work can fail for reasons that pass, like a timeout |
backoff | none | With attempts; exponential for flaky services |
jobId | random | The job must not run twice for the same thing |
delay | 0 | Reminders, grace periods, debouncing |
removeOnComplete | keep all | Always, or Redis grows forever |
Workers in production
Concurrency, rate limits, progress and a clean shutdown. The four settings every worker needs before it goes live.
Jobs one worker runs at once. Raise it for IO bound work.
See the exampleAt most max jobs per duration across all workers of the queue.
See the exampleReport progress; QueueEvents and dashboards see it.
See the exampleStops taking jobs and waits for active ones to finish.
See the exampleconst worker = new Worker('emails', async (job) => {
const batch = chunk(job.data.recipients, 500);
for (const [i, part] of batch.entries()) {
await mailer.send(part);
await job.updateProgress(Math.round(((i + 1) / batch.length) * 100));
}
}, {
connection,
concurrency: 5,
limiter: { max: 50, duration: 1_000 }, // the provider allows 50 calls a second
});
worker.on('failed', (job, err) => log.error({ jobId: job?.id, err }, 'email job failed'));
process.on('SIGTERM', async () => { await worker.close(); process.exit(0); });
Why it matters: a deploy never cuts a job in half, and the mail provider never sees more than it allows.
- SIGTERMdeploy starts
- worker.close()no new jobs taken
- Active jobs finishup to your grace period
- exit(0)the pod stops cleanly
Flows and schedules
Flows run a parent after its children. Job schedulers add a job on a timer or a cron pattern, once per schedule, whatever the number of pods.
Picture it: a parent waits for its children
await new FlowProducer({ connection }).add({
name: 'send-campaign', queueName: 'campaigns', data: { campaignId },
children: [
{ name: 'render-template', queueName: 'templates', data: { campaignId } },
{ name: 'resolve-audience', queueName: 'audience', data: { campaignId } },
{ name: 'check-quota', queueName: 'quota', data: { campaignId } },
],
});
// in the parent worker: const results = await job.getChildrenValues();
Why it matters: each step retries on its own, and the send starts only when everything it needs is ready.
await reports.upsertJobScheduler('daily-grievance-report',
{ pattern: '0 9 * * *', tz: 'Asia/Kolkata' }, // 9am IST every day
{ name: 'grievance-report', data: { recipients: 'management' } });
await sync.upsertJobScheduler('provider-sync', { every: 60_000 }); // every minute
Why it matters: upsert is idempotent, so every pod can call it on boot and there is still one schedule.
BullMQ in NestJS
@nestjs/bullmq registers queues as providers and turns a class into a worker, with dependency injection everywhere.
@Module({
imports: [
BullModule.forRoot({ connection: { url: process.env.REDIS_URL, maxRetriesPerRequest: null } }),
BullModule.registerQueue({ name: 'power-ups', defaultJobOptions: { attempts: 3, removeOnComplete: 1_000 } }),
],
providers: [PowerUpsService, PowerUpsProcessor],
})
export class PowerUpsModule {}
@Injectable()
export class PowerUpsService {
constructor(@InjectQueue('power-ups') private readonly queue: Queue) {}
use(gameId: string, kind: 'magic-word' | 'highlight') { return this.queue.add(kind, { gameId }); }
}
@Processor('power-ups', { concurrency: 10 })
export class PowerUpsProcessor extends WorkerHost {
async process(job: Job<{ gameId: string }>) {
return job.name === 'magic-word' ? this.hints.magicWord(job.data.gameId) : this.hints.highlight(job.data.gameId);
}
@OnWorkerEvent('failed')
onFailed(job: Job, err: Error) { this.log.error({ jobId: job.id, err }); }
}
Why it matters: queues are injected like any provider, and moving slow work behind them is what made the power up APIs about 35% faster.
Production checklist
Redis settings, connections and cleanup that decide whether a queue survives its first busy week.
| Setting | Value | Why |
|---|---|---|
Redis maxmemory-policy | noeviction | Any other policy can silently delete job keys |
Worker maxRetriesPerRequest | null | Blocking commands must wait, not fail after retries |
removeOnComplete and removeOnFail | an age or a count | Finished jobs otherwise stay in Redis forever |
lockDuration | 30000 ms default | Raise it for jobs that block the event loop for long |
| Graceful shutdown | await worker.close() | Deploys do not cut jobs in half |
| Metrics | waiting, active, failed, duration | A growing waiting count is the first sign of trouble |
$ redis-cli CONFIG GET maxmemory-policy 1) "maxmemory-policy" 2) "noeviction" # anything else here and jobs can vanish under memory pressure
Which one do I need?
Name the behaviour you need, then the option or class that gives it to you.
| I want to | Use |
|---|---|
| Answer the API now, do the work later | queue.add() and a Worker |
| Retry flaky work | attempts with backoff: exponential |
| Never run the same job twice | A stable jobId |
| Stay under a provider’s rate limit | Worker limiter: { max, duration } |
| Run something every morning | upsertJobScheduler with a cron pattern |
| Start a step after several others | A flow with FlowProducer |
| Deploy without losing jobs | await worker.close() on SIGTERM |