BullMQ Toolbox 0/7

Backend cheatsheetBullMQ queuesSlow work, off the request

BullMQ
Toolbox

Queues, workers and every state a job passes through, then the options, limits and Redis settings that keep background work reliable once real traffic arrives.

RedisNestJS7 min read
01

Queue, worker, job

A queue is a name in Redis. Producers add jobs to it, workers take them out and run them. Nothing else is needed to start.

BasicsQueueWorkerRedis

Picture it: producers, Redis, workers

APIqueue.add()SchedulerupsertJobSchedulerRedislists, sorted setsWorker 1concurrency 10Worker 2concurrency 10Worker 3concurrency 10QueueEventscompleted, failed
power-ups.tsTypeScript
const connection = new IORedis(process.env.REDIS_URL!, { maxRetriesPerRequest: null });

export const powerUps = new Queue('power-ups', { connection });
await powerUps.add('magic-word', { gameId: 'g42', userId: 'u_91' });

new Worker('power-ups', async (job) => {
  const hint = await words.findHint(job.data.gameId);
  return { hint }; // stored as job.returnvalue
}, { connection, concurrency: 10 });

Why it matters: the API answers at once, and slow work runs on workers you can scale on their own.

Queuenew Queue("power-ups")Adds jobs and inspects the queue. Cheap, use it anywhere.
Workernew Worker("power-ups", fn)Pulls and runs jobs. Holds a blocking Redis connection.
02

The life of a job

Every job moves through a small set of states. Knowing them is how you read the dashboard and debug a stuck queue.

StatesMust knowRetriesStalled

Picture it: every state a job can be in

queue.add()waitingdelayedprioritizedactivecompletedfaileddefaultdelayprioritytimerdoneattempts used upthrew, retry with backoff

Not drawn: a flow parent waits in waiting-children until its children finish, job.retry() sends a failed job back to waiting, and a job whose worker died goes back to waiting once its lock expires.

StateMeansCheck with
waitingReady, no worker free yetgetWaitingCount()
delayedWaiting for a time: a delay or a backoffgetDelayedCount()
prioritizedReady, ordered by prioritygetPrioritizedCount()
activeA worker holds its lockgetActiveCount()
completedReturned a valuegetCompletedCount()
failedThrew and has no attempts leftgetFailed()
waiting-childrenA flow parent waiting for its childrengetWaitingChildrenCount()
03

Adding jobs

Options decide when a job runs, how often it retries, whether duplicates are allowed and how long it is kept.

OptionsattemptsbackoffjobId
add-with-options.tsTypeScript
await emails.add('grievance-report', { date: '2026-09-30' }, {
  jobId: 'grievance-report:2026-09-30',    // same id twice: the second add is ignored
  attempts: 5,
  backoff: { type: 'exponential', delay: 2_000 },
  delay: 60_000,                            // not before one minute from now
  priority: 1,                              // 1 is the highest; omit for plain FIFO
  removeOnComplete: { age: 86_400, count: 1_000 },
  removeOnFail: { age: 7 * 86_400 },
});

await emails.addBulk(recipients.map((r) => ({ name: 'send', data: r }))); // one round trip

Why it matters: a stable jobId makes a nightly job safe to trigger twice, and removeOn options stop Redis from filling up.

How long does a failing job wait between tries? Exponential backoff doubles the delay each time. Run it to see the schedule for the job above.

backoff.jsJavaScript
const delay = 2000;
const attempts = 5;
const exponential = (made) => Math.round(Math.pow(2, made - 1) * delay);

let total = 0;
for (let made = 1; made < attempts; made++) {
  const wait = exponential(made);
  total += wait;
  console.log('after try ' + made + ' wait ' + wait / 1000 + 's');
}
console.log('gives up after ' + attempts + ' tries, ' + total / 1000 + 's of waiting');
TerminalOutput
$ node backoff.js
after try 1 wait 2s
after try 2 wait 4s
after try 3 wait 8s
after try 4 wait 16s
gives up after 5 tries, 30s of waiting

Why it matters: five attempts at a 2 second base spread retries over half a minute, enough to ride out a short outage.

OptionDefaultSet it when
attempts1The work can fail for reasons that pass, like a timeout
backoffnoneWith attempts; exponential for flaky services
jobIdrandomThe job must not run twice for the same thing
delay0Reminders, grace periods, debouncing
removeOnCompletekeep allAlways, or Redis grows forever
04

Workers in production

Concurrency, rate limits, progress and a clean shutdown. The four settings every worker needs before it goes live.

WorkersconcurrencylimiterGraceful shutdown
concurrency: 10

Jobs one worker runs at once. Raise it for IO bound work.

See the example
Throughput
limiter{ max, duration }

At most max jobs per duration across all workers of the queue.

See the example
Rate limit
updateProgress(50)

Report progress; QueueEvents and dashboards see it.

See the example
Feedback
close()

Stops taking jobs and waits for active ones to finish.

See the example
Shutdown
email.worker.tsTypeScript
const worker = new Worker('emails', async (job) => {
  const batch = chunk(job.data.recipients, 500);
  for (const [i, part] of batch.entries()) {
    await mailer.send(part);
    await job.updateProgress(Math.round(((i + 1) / batch.length) * 100));
  }
}, {
  connection,
  concurrency: 5,
  limiter: { max: 50, duration: 1_000 }, // the provider allows 50 calls a second
});

worker.on('failed', (job, err) => log.error({ jobId: job?.id, err }, 'email job failed'));
process.on('SIGTERM', async () => { await worker.close(); process.exit(0); });

Why it matters: a deploy never cuts a job in half, and the mail provider never sees more than it allows.

  1. SIGTERMdeploy starts
  2. worker.close()no new jobs taken
  3. Active jobs finishup to your grace period
  4. exit(0)the pod stops cleanly
05

Flows and schedules

Flows run a parent after its children. Job schedulers add a job on a timer or a cron pattern, once per schedule, whatever the number of pods.

AdvancedFlowProducerupsertJobSchedulerCron

Picture it: a parent waits for its children

send-campaignparent, waitsrender-templatechildresolve-audiencechildcheck-quotachildparent runs when all children finish
campaign.flow.tsTypeScript
await new FlowProducer({ connection }).add({
  name: 'send-campaign', queueName: 'campaigns', data: { campaignId },
  children: [
    { name: 'render-template', queueName: 'templates', data: { campaignId } },
    { name: 'resolve-audience', queueName: 'audience', data: { campaignId } },
    { name: 'check-quota', queueName: 'quota', data: { campaignId } },
  ],
});
// in the parent worker: const results = await job.getChildrenValues();

Why it matters: each step retries on its own, and the send starts only when everything it needs is ready.

schedules.tsTypeScript
await reports.upsertJobScheduler('daily-grievance-report',
  { pattern: '0 9 * * *', tz: 'Asia/Kolkata' },   // 9am IST every day
  { name: 'grievance-report', data: { recipients: 'management' } });

await sync.upsertJobScheduler('provider-sync', { every: 60_000 }); // every minute

Why it matters: upsert is idempotent, so every pod can call it on boot and there is still one schedule.

06

BullMQ in NestJS

@nestjs/bullmq registers queues as providers and turns a class into a worker, with dependency injection everywhere.

Framework@nestjs/bullmqWorkerHost@InjectQueue
power-ups.module.tsTypeScript
@Module({
  imports: [
    BullModule.forRoot({ connection: { url: process.env.REDIS_URL, maxRetriesPerRequest: null } }),
    BullModule.registerQueue({ name: 'power-ups', defaultJobOptions: { attempts: 3, removeOnComplete: 1_000 } }),
  ],
  providers: [PowerUpsService, PowerUpsProcessor],
})
export class PowerUpsModule {}

@Injectable()
export class PowerUpsService {
  constructor(@InjectQueue('power-ups') private readonly queue: Queue) {}
  use(gameId: string, kind: 'magic-word' | 'highlight') { return this.queue.add(kind, { gameId }); }
}

@Processor('power-ups', { concurrency: 10 })
export class PowerUpsProcessor extends WorkerHost {
  async process(job: Job<{ gameId: string }>) {
    return job.name === 'magic-word' ? this.hints.magicWord(job.data.gameId) : this.hints.highlight(job.data.gameId);
  }

  @OnWorkerEvent('failed')
  onFailed(job: Job, err: Error) { this.log.error({ jobId: job.id, err }); }
}

Why it matters: queues are injected like any provider, and moving slow work behind them is what made the power up APIs about 35% faster.

07

Production checklist

Redis settings, connections and cleanup that decide whether a queue survives its first busy week.

OperationsnoevictionmaxRetriesPerRequestMetrics
SettingValueWhy
Redis maxmemory-policynoevictionAny other policy can silently delete job keys
Worker maxRetriesPerRequestnullBlocking commands must wait, not fail after retries
removeOnComplete and removeOnFailan age or a countFinished jobs otherwise stay in Redis forever
lockDuration30000 ms defaultRaise it for jobs that block the event loop for long
Graceful shutdownawait worker.close()Deploys do not cut jobs in half
Metricswaiting, active, failed, durationA growing waiting count is the first sign of trouble
redis-cliTerminal
$ redis-cli CONFIG GET maxmemory-policy
1) "maxmemory-policy"
2) "noeviction"
# anything else here and jobs can vanish under memory pressure

Which one do I need?

Name the behaviour you need, then the option or class that gives it to you.

I want toUse
Answer the API now, do the work laterqueue.add() and a Worker
Retry flaky workattempts with backoff: exponential
Never run the same job twiceA stable jobId
Stay under a provider’s rate limitWorker limiter: { max, duration }
Run something every morningupsertJobScheduler with a cron pattern
Start a step after several othersA flow with FlowProducer
Deploy without losing jobsawait worker.close() on SIGTERM