Backend Engineering GuidesRate limiting
Modules 0/9

Backend Engineering Guides · Rate limiting

Not so
fast.

Rate limiting, explained so anyone can follow it.

Why every API needs a speed limit, how to count the right caller, and how to build it with NestJS and Redis for OTP, emails and payments.

Written byShree Kumar Sharma
9short modules
25 minreading time
1live demo
NestJSwith Redis
REQUESTS / MINUTE LIMIT 5 / MIN 429 slow down YOUR API
SCROLL

01Why it is mandatory

The turnstile every API needs

One button can be pressed a thousand times. A limit makes sure the thousandth press never hurts anyone else.

FundamentalsHTTP 429Retry-AfterMust know
limit SERVER CAPACITY 429

Think of a metro turnstile

The platform can hold only so many people. The turnstile does not care who you are or where you are going. It simply lets a fixed number through each minute, so the platform never overflows, even at rush hour. A rate limiter is that turnstile for one endpoint of your API.

Without one, a single button can be pressed thousands of times. Sometimes that is a buggy app retrying in a loop, sometimes it is a script run on purpose. Either way the server does all the work for every press, and everyone else waits in the same queue.

Four reasons it is not optional

What a refused call looks like

A limited request is not an error in your code. It is a normal, expected answer with status 429 Too Many Requests and a Retry-After header that says how many seconds to wait.

TERMINALbash
$ curl -i -X POST http://localhost:3000/auth/otp -d phone=+919800000001
HTTP/1.1 429 Too Many Requests
Retry-After: 42

{"statusCode":429,"message":"ThrottlerException: Too Many Requests"}

Five words used on every page of this guide

LimitHow many calls are allowed. Example: 5.
WindowThe stretch of time the limit covers. Example: 10 minutes.
KeyWhose calls we are counting. Example: this phone number.
CounterThe running total for one key inside one window.
429The status code a caller gets once the counter passes the limit.

02The limiting window

Clocks, windows and coin jars

Every limit runs on a clock. How that clock resets is the only real difference between the three algorithms.

WindowsFixedSlidingToken bucketMust know
fixed resets sliding bucket +1 every 10 s 1 per request

Your mobile data pack is a fixed window

1.5 GB per day means: the limit is 1.5 GB, the window is one day, and the counter resets at midnight no matter when you started using it. That single sentence holds every idea on this page.

Real windows, and why each one was chosen

ActionLimitWindowWhy this window
Send OTP310 min, fixedAn OTP lives about 5 minutes, so 10 minutes covers a couple of honest resends.
Verify OTP5 triesUntil the code expiresTied to the code, not the clock. A new code starts a new count.
Wrong password5 failures15 min, slidingStops guessing without locking someone out for a typo forever.
Password reset email31 hour, fixedNobody needs more; more usually means someone is flooding an inbox.
Payment attempt31 min per userMatches how fast a person can really retry at a checkout.
Search API60Token bucket, refills 1 per secondTyping sends bursts; the bucket absorbs them.
Free plan exports100Calendar monthA billing period. Lives in the database, see modules 07 and 08.

Interactive

Try a fixed window yourself

Each click runs the same steps as the hitLimit function in module 05, against a pretend Redis key. Click quickly and watch the sixth request get refused.

0of 5 allowed
Window resets inno key yet
redis> GET rl:otp:ip:203.0.113.7 → (nil)
  • Responses will appear here.

The one weak spot of a fixed window

Because the counter resets on the clock, someone can spend the full budget in the last seconds of one window and again in the first seconds of the next. With a limit of 5 per minute, that is 10 calls in about 18 seconds.

Limit: 5 per minute
Window 1: 0 to 60 s Window 2: 60 to 120 s count 5 of 5 count resets, 5 of 5 again 10 accepted in about 18 s 0 s30 s60 s90 s120 s

For OTP and payment limits this rarely matters in practice. When it does, move that one rule to a sliding window.

03Before and after login

Gate or badge?

Before login we only know the door you came through. After login we know your badge. Count the right one.

IdentityPre authPost authtrust proxyMust know
BEFORE LOGIN AFTER LOGIN ip 203.0.113.7 user:42 login swaps the key

Think of an office building

Before you have a badge, the guard can only note which gate you came through and whose floor you asked for. Once you have a badge, the guard just reads the badge number, whichever gate you use. Before login your API is the guard at the gate. After login it reads the badge.

Never count an id the client simply sends

A user id in the request body or a custom header can be anything. Take the user id only from the token your auth guard already verified, and the IP only from the proxy you trust.

One small function decides the key

caller-key.tsTypeScript
export function callerKey(req: Request, action: string): string {
  const user = req.user as { id: string } | undefined; // set by your auth guard
  if (user) return `rl:${action}:user:${user.id}`;    // after login: the badge
  return `rl:${action}:ip:${req.ip}`;                  // before login: the gate
}
Why it matters: the key names both who is calling and what they are doing, so a busy search box never eats the budget meant for payments.

Behind a load balancer every request seems to come from the balancer itself. Tell Express to trust one proxy hop so req.ip becomes the real client address.

main.tsTypeScript
const app = await NestFactory.create<NestExpressApplication>(AppModule);
app.set('trust proxy', 1); // one load balancer in front of us
await app.listen(3000);
TERMINALbash
$ redis-cli --scan --pattern 'rl:*'
rl:otp:ip:203.0.113.7
rl:otp:phone:+919800000001
rl:pay:user:42
# who plus what, readable at a glance
EndpointCaller stateKey onWatch out for
/auth/otpNot logged inPhone number and IPOffices sharing one IP
/auth/loginNot logged inEmail and IPLocking out the real owner
/ordersLogged inUser idTwo devices of the same person
/paymentsLogged inUser id, plus gateway wide capRetries of the same payment
/partner/v1/*Server to serverAPI keyKeys shared across a partner team

04NestJS Throttler

Bolt on a turnstile

Give it two numbers, how many and how long, and NestJS guards every route for you.

NestJS@nestjs/throttlerGuards
@Throttle @SkipThrottle /otp /orders /health

A turnstile you bolt on, not one you build

You tell Throttler two numbers, how many and how long, and it guards every route. You only write code for the few places where the default rule is not right.

TERMINALbash
$ npm i @nestjs/throttler
added 1 package in 2s

Register two named rules once. A short rule stops rapid bursts; a long rule caps steady use. Times are in milliseconds.

app.module.tsTypeScript
@Module({
  imports: [
    ThrottlerModule.forRoot([
      { name: 'short', ttl: 1_000, limit: 3 },   // 3 calls per second
      { name: 'long', ttl: 60_000, limit: 100 }, // 100 calls per minute
    ]),
  ],
  providers: [{ provide: APP_GUARD, useClass: ThrottlerGuard }],
})
export class AppModule {}
What this gives you: every route in the app is limited from now on, with no code in the controllers.

Tighten a sensitive route, or switch the limiter off where it makes no sense, with a decorator.

auth.controller.tsTypeScript
@Controller('auth')
export class AuthController {
  @Post('otp')
  @Throttle({ short: { limit: 1, ttl: 1_000 }, long: { limit: 3, ttl: 600_000 } })
  sendOtp(@Body() dto: SendOtpDto) {
    return this.otp.send(dto.phone);
  }

  @Get('health')
  @SkipThrottle() // load balancers ping this every few seconds
  health() {
    return 'ok';
  }
}
TERMINALbash
$ for i in 1 2 3 4; do curl -s -o /dev/null -w "%{http_code}\n" -X POST localhost:3000/auth/otp; done
201
429
429
429
# the short rule allows one OTP request per second

Teach Throttler who the caller is

By default Throttler counts by IP. Override one method so logged in users are counted by their id, using the same idea as module 03.

caller-throttler.guard.tsTypeScript
@Injectable()
export class CallerThrottlerGuard extends ThrottlerGuard {
  protected async getTracker(req: Record<string, any>): Promise<string> {
    return req.user?.id ? `user:${req.user.id}` : `ip:${req.ip}`;
  }
}
Order matters: authenticate first, then count

A global guard runs before the guards on a controller. If your JWT guard sits on the controller, req.user is still empty when a global Throttler runs, and everyone is counted by IP. Either register the auth guard globally before the throttler, or use @UseGuards(JwtAuthGuard, CallerThrottlerGuard) in that order.

05Redis backed limits

One whiteboard for every cashier

Three servers with three notebooks let one caller in three times. Give them a single shared board.

RedisINCREXPIREioredisMust know
api 1api 2api 3 REDIS INCR rl:user:42 one shared count

Three cashiers, one whiteboard

If each of three cashiers keeps their own notebook, a customer can buy the "limit one per person" item three times by visiting each counter. Put one whiteboard behind all three and the count is honest. With three instances and in memory counting, a limit of 5 quietly becomes 15. Redis is the whiteboard.

TERMINALbash
$ npm i @nest-lab/throttler-storage-redis ioredis
added 3 packages in 3s
app.module.tsTypeScript
ThrottlerModule.forRootAsync({
  inject: [ConfigService],
  useFactory: (config: ConfigService) => ({
    throttlers: [{ ttl: 60_000, limit: 100 }],
    storage: new ThrottlerStorageRedisService(new Redis(config.get('REDIS_URL'))),
  }),
}),
One change, same rules: the decorators from module 04 keep working, but now every instance reads and writes the same counters.

Your own limiter in eight lines

Throttler covers routes. For business rules such as "3 OTPs per phone number", write one small function and reuse it everywhere. It is exactly what the live demo in module 02 does.

rate-limit.tsTypeScript
export async function hitLimit(redis: Redis, key: string, limit: number, windowSec: number) {
  const res = await redis.multi().incr(key).expire(key, windowSec, 'NX').exec();
  const count = Number(res?.[0]?.[1]);
  if (count <= limit) return { allowed: true, retryAfter: 0 };
  return { allowed: false, retryAfter: await redis.ttl(key) };
}

// true only for the first caller inside the time given
export async function claimOnce(redis: Redis, key: string, seconds: number) {
  return (await redis.set(key, '1', 'EX', seconds, 'NX')) === 'OK';
}
How it works: INCR adds one and creates the key at 1 if it is missing. EXPIRE ... NX starts the timer only the first time (Redis 7 and later). MULTI sends both together, so no key is ever left without a timer.
REDIS CLIredis-cli
redis> INCR rl:otp:phone:+919800000001
(integer) 1
redis> EXPIRE rl:otp:phone:+919800000001 600 NX
(integer) 1
redis> INCR rl:otp:phone:+919800000001
(integer) 2
redis> TTL rl:otp:phone:+919800000001
(integer) 583
# 583 seconds until this counter disappears on its own
too-many.tsTypeScript
export function tooMany(retryAfter: number, message = 'Too many requests') {
  return new HttpException({ message, retryAfter }, HttpStatus.TOO_MANY_REQUESTS);
}
Tip: also set the Retry-After header in an exception filter, so apps can wait the right amount instead of guessing.
Need a sliding window instead?

Store each request time in a Redis sorted set, remove entries older than the window, then count what is left. It costs one entry per request instead of one number per window, so keep it for the few rules where fairness at the boundary really matters, such as login failures.

06OTP and emails

Lockers and patient post offices

Refuse the person who asks too often. Queue the email the provider cannot take yet.

OTPEmailBullMQLayered rulesMust know
YOUR CODE 481902 TRIES LEFT RESEND 60s queue at 10 per second EMAILS WAIT, NOT FAIL

A bank locker

You may ask the guard for a new key only once a minute, only a few times an hour, and after three wrong codes the locker locks itself. None of those rules is clever on its own. Together they make the locker safe.

Four rules for OTP

RuleKeyLimitIts one job
Resend cooldownotp:cool:{phone}1 per 60 sStops someone tapping Resend over and over
Per phone numberotp:phone:{phone}3 per 10 minProtects the owner of that number from SMS bombing
Per IPotp:ip:{ip}20 per hourStops one script walking through many numbers
Verify triesotp:try:{phone}5 per codeMakes guessing a six digit code hopeless
1. cooldownOne request per 60 seconds
→
2. per phone3 in 10 minutes
→
3. per IP20 in an hour
→
4. send SMSOnly now do we pay
otp.service.tsTypeScript
async send(phone: string, ip: string) {
  if (!(await claimOnce(this.redis, `otp:cool:${phone}`, 60))) throw tooMany(60);

  const byPhone = await hitLimit(this.redis, `otp:phone:${phone}`, 3, 600);
  if (!byPhone.allowed) throw tooMany(byPhone.retryAfter);

  const byIp = await hitLimit(this.redis, `otp:ip:${ip}`, 20, 3600);
  if (!byIp.allowed) throw tooMany(byIp.retryAfter);

  const code = randomInt(100_000, 1_000_000).toString();
  // store a hash of the code, never the code itself
  await this.redis.set(`otp:code:${phone}`, sha256(code), 'EX', 300);
  await this.sms.send(phone, `Your login code is ${code}`);
}
Read it top to bottom: cheapest check first, most expensive action (the paid SMS) last. If any rule says no, nothing is sent and nothing is billed.
otp.service.tsTypeScript
async verify(phone: string, code: string) {
  const tries = await hitLimit(this.redis, `otp:try:${phone}`, 5, 300);
  if (!tries.allowed) {
    await this.redis.del(`otp:code:${phone}`); // burn the code
    throw tooMany(0, 'Too many wrong codes. Ask for a new one.');
  }
  if ((await this.redis.get(`otp:code:${phone}`)) !== sha256(code)) {
    throw new UnauthorizedException('Wrong code');
  }
  await this.redis.del(`otp:code:${phone}`, `otp:try:${phone}`); // success clears both
}

Emails: say no to the user, but wait for the provider

A post office counter

The counter can serve ten people a minute. It does not send the eleventh person home. It hands them a queue token and serves them a little later. Your email provider is that counter.

Emails have two different limits. The first is what one user may trigger, for example 3 password reset emails per hour. Break it and you answer 429. The second is what your provider can accept, for example 10 emails per second for your account. Break that and you should not refuse anyone; you should slow down. A queue does the slowing for you.

TERMINALbash
$ npm i bullmq
added 1 package in 2s
email.worker.tsTypeScript
export const emailWorker = new Worker('email', (job) => mailer.send(job.data), {
  connection,
  limiter: { max: 10, duration: 1_000 }, // the provider accepts 10 emails per second
});
password-reset.service.tsTypeScript
async requestReset(email: string) {
  const r = await hitLimit(this.redis, `rl:reset:${email}`, 3, 3600); // user rule
  if (!r.allowed) throw tooMany(r.retryAfter);

  await this.emailQueue.add('reset', { to: email }, {
    attempts: 5,
    backoff: { type: 'exponential', delay: 2_000 }, // provider busy? retry later
  });
}
The rule to remember: reject what a user over asks for, queue what a partner cannot take yet.

07Counting after the call

Refunds at the gate

Count the attempt, hand it back when the failure was ours, and let the database guard anything that costs money.

IncrementDecrementIdempotencyTransactionsMust know
1 504 DECR back gateway failed, attempt returned

A prepaid metro card

The gate takes the fare when you tap in. If the gate jams and you never get through, the fare goes back on your card. You pay for rides you took, not for the gate breaking.

Pattern one: refund the attempt when the gateway fails

payment.service.tsTypeScript
async pay(userId: string, dto: PayDto) {
  const key = `rl:pay:user:${userId}`;
  const r = await hitLimit(this.redis, key, 3, 60);
  if (!r.allowed) throw tooMany(r.retryAfter);

  try {
    return await this.gateway.charge(dto);
  } catch (err) {
    // a timeout or 5xx is not the user's fault, so give the attempt back
    if (isServerSideFailure(err)) await this.redis.decr(key);
    throw err;
  }
}
Why only server side failures: a declined card is a real attempt and should count. A gateway timeout is not, and punishing the user for it feels broken.
INCR → 1Attempt counted before the call
→
gateway 504Timeout on their side
→
DECR → 0Attempt handed back

Pattern two: count only wrong passwords

auth.service.tsTypeScript
async login(email: string, password: string) {
  const key = `rl:login:fail:${email}`;
  if (Number(await this.redis.get(key)) >= 5) throw tooMany(await this.redis.ttl(key));

  const user = await this.users.checkPassword(email, password);
  if (!user) {
    await hitLimit(this.redis, key, 5, 900); // one more mistake, window of 15 minutes
    throw new UnauthorizedException('Wrong email or password');
  }
  await this.redis.del(key); // a correct login wipes the slate
  return this.tokens.issue(user);
}

Retries of the same payment must not count twice

Mobile networks drop. The app then sends the same payment again, with the same Idempotency-Key header. The first time we see that key we count and charge. Every repeat just returns the earlier result.

payment.controller.tsTypeScript
@Post('payments')
async create(
  @Req() req: Request,
  @Body() dto: PayDto,
  @Headers('idempotency-key') idemKey: string,
) {
  const firstTime = await claimOnce(this.redis, `idem:${idemKey}`, 86_400);
  // a retry: no new count, no new charge
  if (!firstTime) return this.payments.findByIdemKey(idemKey);
  return this.payments.pay(req.user.id, { ...dto, idemKey });
}

Pattern three: money limits live in the database

A free transfer quota decides whether someone is charged. If the transfer fails, the quota must come back automatically, and it must survive a restart. That is exactly what a database transaction promises and what a Redis counter does not.

transfer.service.tsTypeScript
async transfer(userId: string, dto: TransferDto) {
  return this.db.transaction(async (tx) => {
    const used = await tx.query(
      `UPDATE wallets SET free_transfers = free_transfers - 1
        WHERE user_id = $1 AND free_transfers > 0`, [userId]);
    if (used.rowCount === 0) throw new HttpException('No free transfers left', 402);

    // if this insert fails, the decrement above rolls back with it
    return tx.query(
      `INSERT INTO transfers (user_id, amount, to_account) VALUES ($1, $2, $3)`,
      [userId, dto.amount, dto.toAccount],
    );
  });
}
Why it is safe: the WHERE free_transfers > 0 check and the decrement happen in one statement, so two requests at the same moment can never both spend the last free transfer.
Redis for speed limits, the database for money limits

Speed limits are short lived and losing one on a restart costs nothing. Quotas tied to billing must be exact and must roll back with the payment, so keep them in the same database transaction as the thing they protect.

08Cron or Redis expiry

Alarm clocks and expiry dates

Some counters throw themselves away. Others need a clock to reset them. Scale decides which.

CronTTL@nestjs/scheduleScale
0 0 * * * KEY WITH A TTL rl:otp:phone deletes itself at 0 EXPIRE 600

Reading a cron expression

A cron expression is five fields separated by spaces. A star means "every". Hover a box to see which part of the expression it controls.

0minute
0 to 59
0hour
0 to 23
1day of month
1 to 31
*month
every month
*day of week
any day

0 0 1 * * means: at minute 0 of hour 0, on day 1, of every month. In plain words, midnight on the first of each month. 0 0 * * * is every midnight; */15 * * * * is every fifteen minutes.

When a limiter really needs cron

Resetting calendar quotas stored in the database, such as 100 exports per month on the free plan. Lifting temporary bans saved as rows. Deleting old login attempt logs. Sending the nightly report of who got limited most. All four are about the calendar or about bulk work, not about the next few seconds.

TERMINALbash
$ npm i @nestjs/schedule
added 2 packages in 2s
quota-reset.job.tsTypeScript
@Injectable()
export class QuotaResetJob {
  constructor(private readonly db: Db, private readonly redis: Redis) {}

  @Cron('0 0 1 * *', { timeZone: 'Asia/Kolkata' }) // midnight on the 1st, Indian time
  async resetMonthlyExports() {
    const gotLock = await claimOnce(this.redis, 'lock:quota-reset', 300);
    if (!gotLock) return; // another instance is already doing it
    await this.db.query(`UPDATE plans SET exports_used = 0`);
  }
}
Two traps it avoids: without the lock, every instance runs the same job at midnight. Without the time zone, "midnight" means the server clock, often UTC.

Often you can skip the job entirely: reset on read

Store the start of the current period next to the counter. When a request arrives in a new month, the same statement resets and counts. No job, no midnight spike, nothing to monitor.

exports.sqlSQL
-- count one export; start from 1 again if a new month has begun
UPDATE plans
   SET exports_used = CASE WHEN period_start < date_trunc('month', now())
                           THEN 1 ELSE exports_used + 1 END,
       period_start = date_trunc('month', now())
 WHERE user_id = $1
   AND (period_start < date_trunc('month', now()) OR exports_used < export_limit)
RETURNING exports_used; -- no row back: limit reached

Which one, by scale and usage

SituationUseWhy
One instance, windows of seconds or minutesThrottler in memoryNothing to share, nothing to reset
Several instances, short windowsRedis keys with a TTLShared by all, cleans itself up, no job to run
Daily or monthly quotas tied to a plan or billDatabase row with reset on readExact, survives restarts, rolls back with transactions
Bulk clean up, reports, lifting saved bansCron with a Redis lockRuns once, at a known time, on one instance
Many instances and many scheduled jobs that need retriesBullMQ job schedulerThe queue decides who runs it and retries failures
Millions of users on one reset timeReset on read, or stagger by userAvoids one huge UPDATE and a traffic spike at 00:00
schedule.tsTypeScript
await queue.upsertJobScheduler('reset-daily-sms', {
  pattern: '0 0 * * *', // every midnight
  tz: 'Asia/Kolkata',
});
When to prefer this over @Cron: the job is stored in Redis, so exactly one worker picks it up, failed runs retry, and you can see its history.

09Side by side

Keep this in your pocket

Every rule from this guide on one page, and a checklist to run before you ship.

SummaryChecklistGlossary
QuestionOTPEmailsPaymentsLogin
Caller stateNot logged inEitherLogged inNot logged in
Key onPhone and IPEmail addressUser idEmail and IP
Typical rule3 per 10 min3 resets per hour3 per minute5 failures per 15 min
When to countBefore sendingBefore queuingBefore, refund on 5xxOnly on failure
Where it livesRedis TTLRedis TTL plus a queueRedis TTL, quotas in the DBRedis TTL
1

Before login, count the gate

IP plus the phone or email being targeted.

2

After login, count the badge

The verified user id, never an id the client sends.

3

More than one server? Redis

One shared counter, with a timer on every key.

4

Money limits go in the database

Inside the same transaction as the payment.

Before you ship a limiter

Glossary

429 Too Many Requests
The answer a caller gets once they pass the limit.
Retry-After
A header with the number of seconds to wait before trying again.
TTL
Time to live. Seconds left before a Redis key deletes itself.
Window
The stretch of time a limit covers, such as 10 minutes.
NAT
Many devices sharing one public IP, like everyone in an office.
Atomic
Done completely or not at all, with nothing else slipping in between.
Idempotency key
A unique id per action, so a retried request is recognised as the same one.
Cron
A schedule that runs a task at fixed times, written as five fields.
Lock
A short lived key that lets only one instance do a job at a time.
Token bucket
A limit that refills slowly, allowing short bursts but a steady average.
Back to the top