01Why it is mandatory
The turnstile every API needs
One button can be pressed a thousand times. A limit makes sure the thousandth press never hurts anyone else.
Think of a metro turnstile
The platform can hold only so many people. The turnstile does not care who you are or where you are going. It simply lets a fixed number through each minute, so the platform never overflows, even at rush hour. A rate limiter is that turnstile for one endpoint of your API.
Without one, a single button can be pressed thousands of times. Sometimes that is a buggy app retrying in a loop, sometimes it is a script run on purpose. Either way the server does all the work for every press, and everyone else waits in the same queue.
Four reasons it is not optional
Some calls cost real money
Every OTP SMS and every email is billed by the provider. A script that requests ten thousand OTPs overnight is a real invoice the next morning.
See the OTP rules SecurityGuessing becomes too slow
Cracking a password or a six digit OTP is a numbers game. Five tries per code turns a million guesses into a game nobody can win.
See the login rule FairnessOne noisy caller cannot slow everyone
CPU, memory and database connections are shared. A limit keeps one runaway client from making the page slow for every real user.
See shared counters PartnersYour providers have limits too
Payment gateways and email services cap how fast you may call them. Go over and they throttle your whole account, not the one user who caused it.
See the email queueWhat a refused call looks like
A limited request is not an error in your code. It is a normal, expected answer with status 429 Too Many Requests and a Retry-After header that says how many seconds to wait.
$ curl -i -X POST http://localhost:3000/auth/otp -d phone=+919800000001 HTTP/1.1 429 Too Many Requests Retry-After: 42 {"statusCode":429,"message":"ThrottlerException: Too Many Requests"}
Five words used on every page of this guide
02The limiting window
Clocks, windows and coin jars
Every limit runs on a clock. How that clock resets is the only real difference between the three algorithms.
Your mobile data pack is a fixed window
1.5 GB per day means: the limit is 1.5 GB, the window is one day, and the counter resets at midnight no matter when you started using it. That single sentence holds every idea on this page.
Resets on the clock
Like the daily data pack. One counter per slice of time, so it is cheap and easy to explain. Weak spot: a burst right at the reset moment.
Start hereTry it live Sliding windowAlways looks back from now
Like a library that allows three books in any fourteen days. It counts what happened in the last N seconds from this moment, so there is no reset moment to exploit.
FairerBuild it in Redis Token bucketA jar that refills slowly
A jar holds five coins and gains one every ten seconds. Each request spends a coin. Short bursts are fine, the long run average stays steady.
Bursty trafficSee real windowsReal windows, and why each one was chosen
| Action | Limit | Window | Why this window |
|---|---|---|---|
| Send OTP | 3 | 10 min, fixed | An OTP lives about 5 minutes, so 10 minutes covers a couple of honest resends. |
| Verify OTP | 5 tries | Until the code expires | Tied to the code, not the clock. A new code starts a new count. |
| Wrong password | 5 failures | 15 min, sliding | Stops guessing without locking someone out for a typo forever. |
| Password reset email | 3 | 1 hour, fixed | Nobody needs more; more usually means someone is flooding an inbox. |
| Payment attempt | 3 | 1 min per user | Matches how fast a person can really retry at a checkout. |
| Search API | 60 | Token bucket, refills 1 per second | Typing sends bursts; the bucket absorbs them. |
| Free plan exports | 100 | Calendar month | A billing period. Lives in the database, see modules 07 and 08. |
Interactive
Try a fixed window yourself
Each click runs the same steps as the hitLimit function in module 05, against a pretend Redis key. Click quickly and watch the sixth request get refused.
- Responses will appear here.
The one weak spot of a fixed window
Because the counter resets on the clock, someone can spend the full budget in the last seconds of one window and again in the first seconds of the next. With a limit of 5 per minute, that is 10 calls in about 18 seconds.
For OTP and payment limits this rarely matters in practice. When it does, move that one rule to a sliding window.
03Before and after login
Gate or badge?
Before login we only know the door you came through. After login we know your badge. Count the right one.
Think of an office building
Before you have a badge, the guard can only note which gate you came through and whose floor you asked for. Once you have a badge, the guard just reads the badge number, whichever gate you use. Before login your API is the guard at the gate. After login it reads the badge.
Count the gate and the target
IP address is always present, but many people can share one: an office, a college, a mobile network. The target, the phone number or email being asked for, protects the person on the receiving end even when the attacker changes IPs. A device id, a random id your app stores on first launch, helps too, though it can be cleared.
See it in NestJS After loginCount the badge
The user id from the verified token is the best key there is: stable, unique and the same on every device. Add the session or device when each device should get its own budget. Servers calling you use an API key; company plans use the organisation id.
See it in NestJSA user id in the request body or a custom header can be anything. Take the user id only from the token your auth guard already verified, and the IP only from the proxy you trust.
One small function decides the key
export function callerKey(req: Request, action: string): string { const user = req.user as { id: string } | undefined; // set by your auth guard if (user) return `rl:${action}:user:${user.id}`; // after login: the badge return `rl:${action}:ip:${req.ip}`; // before login: the gate }
Behind a load balancer every request seems to come from the balancer itself. Tell Express to trust one proxy hop so req.ip becomes the real client address.
const app = await NestFactory.create<NestExpressApplication>(AppModule); app.set('trust proxy', 1); // one load balancer in front of us await app.listen(3000);
$ redis-cli --scan --pattern 'rl:*' rl:otp:ip:203.0.113.7 rl:otp:phone:+919800000001 rl:pay:user:42 # who plus what, readable at a glance
| Endpoint | Caller state | Key on | Watch out for |
|---|---|---|---|
| /auth/otp | Not logged in | Phone number and IP | Offices sharing one IP |
| /auth/login | Not logged in | Email and IP | Locking out the real owner |
| /orders | Logged in | User id | Two devices of the same person |
| /payments | Logged in | User id, plus gateway wide cap | Retries of the same payment |
| /partner/v1/* | Server to server | API key | Keys shared across a partner team |
04NestJS Throttler
Bolt on a turnstile
Give it two numbers, how many and how long, and NestJS guards every route for you.
A turnstile you bolt on, not one you build
You tell Throttler two numbers, how many and how long, and it guards every route. You only write code for the few places where the default rule is not right.
$ npm i @nestjs/throttler added 1 package in 2s
Register two named rules once. A short rule stops rapid bursts; a long rule caps steady use. Times are in milliseconds.
@Module({ imports: [ ThrottlerModule.forRoot([ { name: 'short', ttl: 1_000, limit: 3 }, // 3 calls per second { name: 'long', ttl: 60_000, limit: 100 }, // 100 calls per minute ]), ], providers: [{ provide: APP_GUARD, useClass: ThrottlerGuard }], }) export class AppModule {}
Tighten a sensitive route, or switch the limiter off where it makes no sense, with a decorator.
@Controller('auth') export class AuthController { @Post('otp') @Throttle({ short: { limit: 1, ttl: 1_000 }, long: { limit: 3, ttl: 600_000 } }) sendOtp(@Body() dto: SendOtpDto) { return this.otp.send(dto.phone); } @Get('health') @SkipThrottle() // load balancers ping this every few seconds health() { return 'ok'; } }
$ for i in 1 2 3 4; do curl -s -o /dev/null -w "%{http_code}\n" -X POST localhost:3000/auth/otp; done 201 429 429 429 # the short rule allows one OTP request per second
Teach Throttler who the caller is
By default Throttler counts by IP. Override one method so logged in users are counted by their id, using the same idea as module 03.
@Injectable() export class CallerThrottlerGuard extends ThrottlerGuard { protected async getTracker(req: Record<string, any>): Promise<string> { return req.user?.id ? `user:${req.user.id}` : `ip:${req.ip}`; } }
A global guard runs before the guards on a controller. If your JWT guard sits on the controller, req.user is still empty when a global Throttler runs, and everyone is counted by IP. Either register the auth guard globally before the throttler, or use @UseGuards(JwtAuthGuard, CallerThrottlerGuard) in that order.
05Redis backed limits
One whiteboard for every cashier
Three servers with three notebooks let one caller in three times. Give them a single shared board.
Three cashiers, one whiteboard
If each of three cashiers keeps their own notebook, a customer can buy the "limit one per person" item three times by visiting each counter. Put one whiteboard behind all three and the count is honest. With three instances and in memory counting, a limit of 5 quietly becomes 15. Redis is the whiteboard.
$ npm i @nest-lab/throttler-storage-redis ioredis added 3 packages in 3s
ThrottlerModule.forRootAsync({ inject: [ConfigService], useFactory: (config: ConfigService) => ({ throttlers: [{ ttl: 60_000, limit: 100 }], storage: new ThrottlerStorageRedisService(new Redis(config.get('REDIS_URL'))), }), }),
Your own limiter in eight lines
Throttler covers routes. For business rules such as "3 OTPs per phone number", write one small function and reuse it everywhere. It is exactly what the live demo in module 02 does.
export async function hitLimit(redis: Redis, key: string, limit: number, windowSec: number) { const res = await redis.multi().incr(key).expire(key, windowSec, 'NX').exec(); const count = Number(res?.[0]?.[1]); if (count <= limit) return { allowed: true, retryAfter: 0 }; return { allowed: false, retryAfter: await redis.ttl(key) }; } // true only for the first caller inside the time given export async function claimOnce(redis: Redis, key: string, seconds: number) { return (await redis.set(key, '1', 'EX', seconds, 'NX')) === 'OK'; }
INCR adds one and creates the key at 1 if it is missing. EXPIRE ... NX starts the timer only the first time (Redis 7 and later). MULTI sends both together, so no key is ever left without a timer.redis> INCR rl:otp:phone:+919800000001 (integer) 1 redis> EXPIRE rl:otp:phone:+919800000001 600 NX (integer) 1 redis> INCR rl:otp:phone:+919800000001 (integer) 2 redis> TTL rl:otp:phone:+919800000001 (integer) 583 # 583 seconds until this counter disappears on its own
export function tooMany(retryAfter: number, message = 'Too many requests') { return new HttpException({ message, retryAfter }, HttpStatus.TOO_MANY_REQUESTS); }
Retry-After header in an exception filter, so apps can wait the right amount instead of guessing.Store each request time in a Redis sorted set, remove entries older than the window, then count what is left. It costs one entry per request instead of one number per window, so keep it for the few rules where fairness at the boundary really matters, such as login failures.
06OTP and emails
Lockers and patient post offices
Refuse the person who asks too often. Queue the email the provider cannot take yet.
A bank locker
You may ask the guard for a new key only once a minute, only a few times an hour, and after three wrong codes the locker locks itself. None of those rules is clever on its own. Together they make the locker safe.
Four rules for OTP
| Rule | Key | Limit | Its one job |
|---|---|---|---|
| Resend cooldown | otp:cool:{phone} | 1 per 60 s | Stops someone tapping Resend over and over |
| Per phone number | otp:phone:{phone} | 3 per 10 min | Protects the owner of that number from SMS bombing |
| Per IP | otp:ip:{ip} | 20 per hour | Stops one script walking through many numbers |
| Verify tries | otp:try:{phone} | 5 per code | Makes guessing a six digit code hopeless |
async send(phone: string, ip: string) { if (!(await claimOnce(this.redis, `otp:cool:${phone}`, 60))) throw tooMany(60); const byPhone = await hitLimit(this.redis, `otp:phone:${phone}`, 3, 600); if (!byPhone.allowed) throw tooMany(byPhone.retryAfter); const byIp = await hitLimit(this.redis, `otp:ip:${ip}`, 20, 3600); if (!byIp.allowed) throw tooMany(byIp.retryAfter); const code = randomInt(100_000, 1_000_000).toString(); // store a hash of the code, never the code itself await this.redis.set(`otp:code:${phone}`, sha256(code), 'EX', 300); await this.sms.send(phone, `Your login code is ${code}`); }
async verify(phone: string, code: string) { const tries = await hitLimit(this.redis, `otp:try:${phone}`, 5, 300); if (!tries.allowed) { await this.redis.del(`otp:code:${phone}`); // burn the code throw tooMany(0, 'Too many wrong codes. Ask for a new one.'); } if ((await this.redis.get(`otp:code:${phone}`)) !== sha256(code)) { throw new UnauthorizedException('Wrong code'); } await this.redis.del(`otp:code:${phone}`, `otp:try:${phone}`); // success clears both }
Emails: say no to the user, but wait for the provider
A post office counter
The counter can serve ten people a minute. It does not send the eleventh person home. It hands them a queue token and serves them a little later. Your email provider is that counter.
Emails have two different limits. The first is what one user may trigger, for example 3 password reset emails per hour. Break it and you answer 429. The second is what your provider can accept, for example 10 emails per second for your account. Break that and you should not refuse anyone; you should slow down. A queue does the slowing for you.
$ npm i bullmq added 1 package in 2s
export const emailWorker = new Worker('email', (job) => mailer.send(job.data), { connection, limiter: { max: 10, duration: 1_000 }, // the provider accepts 10 emails per second });
async requestReset(email: string) { const r = await hitLimit(this.redis, `rl:reset:${email}`, 3, 3600); // user rule if (!r.allowed) throw tooMany(r.retryAfter); await this.emailQueue.add('reset', { to: email }, { attempts: 5, backoff: { type: 'exponential', delay: 2_000 }, // provider busy? retry later }); }
07Counting after the call
Refunds at the gate
Count the attempt, hand it back when the failure was ours, and let the database guard anything that costs money.
A prepaid metro card
The gate takes the fare when you tap in. If the gate jams and you never get through, the fare goes back on your card. You pay for rides you took, not for the gate breaking.
Count first, refund our failures
Add one before calling the payment gateway. If the call fails because of us or the gateway, a timeout or a 5xx, take the one back.
Payments, OTP sendsShow the code Pattern twoCount only mistakes
Add one only when the password or code is wrong. A correct login wipes the counter clean, so honest users are never slowed.
Login, OTP verifyShow the code Pattern threeCount only what succeeded
Spend from a balance only when the action commits, such as "10 free transfers a month". Belongs in the database, inside the same transaction.
Quotas, billingShow the codePattern one: refund the attempt when the gateway fails
async pay(userId: string, dto: PayDto) { const key = `rl:pay:user:${userId}`; const r = await hitLimit(this.redis, key, 3, 60); if (!r.allowed) throw tooMany(r.retryAfter); try { return await this.gateway.charge(dto); } catch (err) { // a timeout or 5xx is not the user's fault, so give the attempt back if (isServerSideFailure(err)) await this.redis.decr(key); throw err; } }
Pattern two: count only wrong passwords
async login(email: string, password: string) { const key = `rl:login:fail:${email}`; if (Number(await this.redis.get(key)) >= 5) throw tooMany(await this.redis.ttl(key)); const user = await this.users.checkPassword(email, password); if (!user) { await hitLimit(this.redis, key, 5, 900); // one more mistake, window of 15 minutes throw new UnauthorizedException('Wrong email or password'); } await this.redis.del(key); // a correct login wipes the slate return this.tokens.issue(user); }
Retries of the same payment must not count twice
Mobile networks drop. The app then sends the same payment again, with the same Idempotency-Key header. The first time we see that key we count and charge. Every repeat just returns the earlier result.
@Post('payments') async create( @Req() req: Request, @Body() dto: PayDto, @Headers('idempotency-key') idemKey: string, ) { const firstTime = await claimOnce(this.redis, `idem:${idemKey}`, 86_400); // a retry: no new count, no new charge if (!firstTime) return this.payments.findByIdemKey(idemKey); return this.payments.pay(req.user.id, { ...dto, idemKey }); }
Pattern three: money limits live in the database
A free transfer quota decides whether someone is charged. If the transfer fails, the quota must come back automatically, and it must survive a restart. That is exactly what a database transaction promises and what a Redis counter does not.
async transfer(userId: string, dto: TransferDto) { return this.db.transaction(async (tx) => { const used = await tx.query( `UPDATE wallets SET free_transfers = free_transfers - 1 WHERE user_id = $1 AND free_transfers > 0`, [userId]); if (used.rowCount === 0) throw new HttpException('No free transfers left', 402); // if this insert fails, the decrement above rolls back with it return tx.query( `INSERT INTO transfers (user_id, amount, to_account) VALUES ($1, $2, $3)`, [userId, dto.amount, dto.toAccount], ); }); }
WHERE free_transfers > 0 check and the decrement happen in one statement, so two requests at the same moment can never both spend the last free transfer.Speed limits are short lived and losing one on a restart costs nothing. Quotas tied to billing must be exact and must roll back with the payment, so keep them in the same database transaction as the thing they protect.
08Cron or Redis expiry
Alarm clocks and expiry dates
Some counters throw themselves away. Others need a clock to reset them. Scale decides which.
A milk carton with an expiry date
The key carries its own timer and disappears when it runs out. No job, no schedule, nothing to forget. Every counter in modules 05 to 07 works like this.
See TTL keys CronAn alarm clock for code
Cron runs a task at fixed times, such as every night at midnight or on the first of each month. It suits calendar limits and big clean up work.
Read an expressionReading a cron expression
A cron expression is five fields separated by spaces. A star means "every". Hover a box to see which part of the expression it controls.
0 to 59
0 to 23
1 to 31
every month
any day
0 0 1 * * means: at minute 0 of hour 0, on day 1, of every month. In plain words, midnight on the first of each month. 0 0 * * * is every midnight; */15 * * * * is every fifteen minutes.
When a limiter really needs cron
Resetting calendar quotas stored in the database, such as 100 exports per month on the free plan. Lifting temporary bans saved as rows. Deleting old login attempt logs. Sending the nightly report of who got limited most. All four are about the calendar or about bulk work, not about the next few seconds.
$ npm i @nestjs/schedule added 2 packages in 2s
@Injectable() export class QuotaResetJob { constructor(private readonly db: Db, private readonly redis: Redis) {} @Cron('0 0 1 * *', { timeZone: 'Asia/Kolkata' }) // midnight on the 1st, Indian time async resetMonthlyExports() { const gotLock = await claimOnce(this.redis, 'lock:quota-reset', 300); if (!gotLock) return; // another instance is already doing it await this.db.query(`UPDATE plans SET exports_used = 0`); } }
Often you can skip the job entirely: reset on read
Store the start of the current period next to the counter. When a request arrives in a new month, the same statement resets and counts. No job, no midnight spike, nothing to monitor.
-- count one export; start from 1 again if a new month has begun UPDATE plans SET exports_used = CASE WHEN period_start < date_trunc('month', now()) THEN 1 ELSE exports_used + 1 END, period_start = date_trunc('month', now()) WHERE user_id = $1 AND (period_start < date_trunc('month', now()) OR exports_used < export_limit) RETURNING exports_used; -- no row back: limit reached
Which one, by scale and usage
| Situation | Use | Why |
|---|---|---|
| One instance, windows of seconds or minutes | Throttler in memory | Nothing to share, nothing to reset |
| Several instances, short windows | Redis keys with a TTL | Shared by all, cleans itself up, no job to run |
| Daily or monthly quotas tied to a plan or bill | Database row with reset on read | Exact, survives restarts, rolls back with transactions |
| Bulk clean up, reports, lifting saved bans | Cron with a Redis lock | Runs once, at a known time, on one instance |
| Many instances and many scheduled jobs that need retries | BullMQ job scheduler | The queue decides who runs it and retries failures |
| Millions of users on one reset time | Reset on read, or stagger by user | Avoids one huge UPDATE and a traffic spike at 00:00 |
await queue.upsertJobScheduler('reset-daily-sms', { pattern: '0 0 * * *', // every midnight tz: 'Asia/Kolkata', });
09Side by side
Keep this in your pocket
Every rule from this guide on one page, and a checklist to run before you ship.
| Question | OTP | Emails | Payments | Login |
|---|---|---|---|---|
| Caller state | Not logged in | Either | Logged in | Not logged in |
| Key on | Phone and IP | Email address | User id | Email and IP |
| Typical rule | 3 per 10 min | 3 resets per hour | 3 per minute | 5 failures per 15 min |
| When to count | Before sending | Before queuing | Before, refund on 5xx | Only on failure |
| Where it lives | Redis TTL | Redis TTL plus a queue | Redis TTL, quotas in the DB | Redis TTL |
Before login, count the gate
IP plus the phone or email being targeted.
After login, count the badge
The verified user id, never an id the client sends.
More than one server? Redis
One shared counter, with a timer on every key.
Money limits go in the database
Inside the same transaction as the payment.
Before you ship a limiter
Glossary
- 429 Too Many Requests
- The answer a caller gets once they pass the limit.
- Retry-After
- A header with the number of seconds to wait before trying again.
- TTL
- Time to live. Seconds left before a Redis key deletes itself.
- Window
- The stretch of time a limit covers, such as 10 minutes.
- NAT
- Many devices sharing one public IP, like everyone in an office.
- Atomic
- Done completely or not at all, with nothing else slipping in between.
- Idempotency key
- A unique id per action, so a retried request is recognised as the same one.
- Cron
- A schedule that runs a task at fixed times, written as five fields.
- Lock
- A short lived key that lets only one instance do a job at a time.
- Token bucket
- A limit that refills slowly, allowing short bursts but a steady average.