API Rate Limiting: Fixed Window, Sliding Log and Token Bucket, and Making Them Atomic in Redis
Key takeaways
Rate limiting protects your API from abuse, ensures fair use, and prevents cascading failures. This guide covers the three main algorithms, Redis-backed implementation, and production patterns for Node.js APIs.
Why Rate Limiting Matters
Without rate limiting:
- A single script can exhaust your database connections
- One angry user can DoS your API
- Credential stuffing attacks run unhindered
- You have no way to enforce pricing tiers
The list above makes rate limiting sound like a security feature, but in practice it is mostly a capacity feature. The traffic that hurts you is rarely an attacker. It is usually a customer’s cron job that someone pointed at your API with no sleep between calls, or a mobile client whose retry logic spins in a tight loop when your backend returns 500s. A limiter’s real job is to keep one misbehaving client from using up the shared resources (database connections, worker threads, a downstream vendor’s quota) that every other client needs.
That framing changes what you optimize for. You rarely need a precise limit. What you need is a limit that is cheap to check on every request, consistent across all your API instances, and predictable for clients, so a well-behaved client can tell how fast it may go and back off correctly. Keep those three properties in mind when you compare the algorithms below.
Rate Limiting Algorithms
Fixed Window
Count requests in a fixed time window (e.g., 100 requests per minute).
Minute 0:00-1:00 → 100 requests allowed
Minute 1:00-2:00 → counter resets, 100 more allowed
Problem: A user can make 100 requests at 0:59 and 100 more at 1:01 — 200 requests in 2 seconds.
// Redis implementation
async function fixedWindowLimit(key, limit, windowSeconds) {
const now = Math.floor(Date.now() / 1000)
const windowKey = `ratelimit:${key}:${Math.floor(now / windowSeconds)}`
const count = await redis.incr(windowKey)
if (count === 1) await redis.expire(windowKey, windowSeconds)
return { allowed: count <= limit, count, limit }
}
There is a subtle bug in this snippet. INCR and EXPIRE are two separate round trips. If the process crashes, or the Redis connection drops, between them, the key is left without a TTL. The windowed key name means that stale key is never read again, so the limit still works, but it sits in Redis forever. At a million distinct users per day that is a slow memory leak you only notice when used_memory alarms go off months later. Fix it with SET key 0 EX <window> NX followed by INCR, with a MULTI block, or with a short Lua script. Every Redis counter pattern needs this, not just rate limiting.
The boundary burst is the reason people move away from fixed windows, but it matters less than it looks. The worst case is 2× the limit over one window length. If your limit is set well below real capacity, 2× for a few seconds is absorbable. Fixed window is still a reasonable choice for coarse, high-volume limits (for example, “10k requests per hour per API key”) because it costs one small key per client.
Sliding Window Log
Track exact timestamps of each request. Count requests in the last N seconds.
async function slidingWindowLog(key, limit, windowSeconds) {
const now = Date.now()
const windowStart = now - windowSeconds * 1000
const logKey = `ratelimit:log:${key}`
const pipeline = redis.pipeline()
pipeline.zremrangebyscore(logKey, 0, windowStart) // remove old entries
pipeline.zadd(logKey, now, `${now}-${Math.random()}`) // add current
pipeline.zcard(logKey) // count in window
pipeline.expire(logKey, windowSeconds)
const results = await pipeline.exec()
const count = results[2][1]
return { allowed: count <= limit, count, limit }
}
Accurate but memory-intensive — stores one entry per request.
Look at the order of the pipeline: the request is added to the log before the check, and it is added whether it is allowed or not. So rejected requests count toward the limit too. A client that ignores 429s and keeps hammering will never see its window drain. It stays locked out for as long as it keeps retrying. Sometimes that is what you want (it punishes abusive clients). Often it is not: a buggy SDK with no backoff can lock a paying customer out indefinitely. If you want only successful requests to count, check ZCARD first and ZADD only when allowed. That takes a Lua script, because a check-then-write across two round trips races between concurrent requests.
Memory is the other cost. A limit of 10,000 requests per hour means up to 10,000 sorted-set members per client, each about 50–70 bytes with the random suffix. For high limits, the sliding window counter approximation is usually a better fit: keep the fixed-window counts for the current and previous windows, and weight the previous one by how much of it still overlaps the sliding window. You get close to sliding-log accuracy with only two integers per client.
Token Bucket (Recommended)
A bucket fills with tokens at a constant rate. Each request consumes one token. Allows bursts up to bucket capacity.
// Token bucket with Redis
async function tokenBucket(key, capacity, refillRate, refillSeconds) {
const now = Date.now() / 1000 // seconds
const bucketKey = `ratelimit:bucket:${key}`
// Lua script for atomicity
const script = `
local key = KEYS[1]
local capacity = tonumber(ARGV[1])
local refill_rate = tonumber(ARGV[2])
local refill_seconds = tonumber(ARGV[3])
local now = tonumber(ARGV[4])
local bucket = redis.call('HMGET', key, 'tokens', 'last_refill')
local tokens = tonumber(bucket[1]) or capacity
local last_refill = tonumber(bucket[2]) or now
-- Refill tokens based on elapsed time
local elapsed = now - last_refill
local new_tokens = math.min(capacity, tokens + (elapsed * refill_rate / refill_seconds))
if new_tokens < 1 then
redis.call('HMSET', key, 'tokens', new_tokens, 'last_refill', now)
redis.call('EXPIRE', key, refill_seconds * 2)
return {0, math.ceil((1 - new_tokens) * refill_seconds / refill_rate)}
end
redis.call('HMSET', key, 'tokens', new_tokens - 1, 'last_refill', now)
redis.call('EXPIRE', key, refill_seconds * 2)
return {1, 0}
`
const [allowed, retryAfter] = await redis.eval(
script, 1, bucketKey, capacity, refillRate, refillSeconds, now
)
return { allowed: allowed === 1, retryAfter }
}
// Usage: 100 requests per minute, burst up to 20
await tokenBucket('user:123', 20, 100, 60)
Two design choices here matter more than the algorithm itself.
Why Lua? The script reads the bucket, computes the refill, and writes it back. If you did that as separate HMGET and HMSET calls from Node, two concurrent requests could both read “1 token left”, both decide they are allowed, and both write back 0. Redis runs a Lua script atomically, so there is no gap between read and write. The cost is that a slow script blocks all of Redis. Keep these scripts tiny and never loop over large keys inside them.
Whose clock? The script uses now from the application server. With several API instances, each instance’s clock is slightly different. If one instance runs 2 seconds fast, its requests “refill” tokens that the other instances think do not exist yet, and when a slower instance writes a smaller last_refill, elapsed goes negative. The more robust option is to call redis.call('TIME') inside the script, so every instance uses Redis’s clock. (On Redis versions older than 5, this requires redis.replicate_commands() first, because scripts were replicated verbatim and a non-deterministic call was rejected.)
This is the bug I would look for first if a token bucket ever “leaks” in production, because it is hard to reproduce locally. The symptom is a limiter that occasionally lets clients through well above the configured rate, only for a few minutes, and often right after a deploy or scale-out. The limiter code has not changed, so you end up suspecting Redis. The usual cause is a newly started instance whose clock is not yet NTP-synced and runs a few seconds off from the rest of the fleet. Its refill calculations hand out extra tokens until the clock settles. With one instance on your laptop you will never see it. Switching the script to Redis TIME removes the whole class of problems, which is why I use it by default.
Express Middleware
Using express-rate-limit (simple)
npm install express-rate-limit rate-limit-redis
import rateLimit from 'express-rate-limit'
import { RedisStore } from 'rate-limit-redis'
import { redisClient } from './redis'
// Global rate limit
const globalLimit = rateLimit({
windowMs: 15 * 60 * 1000, // 15 minutes
max: 1000,
standardHeaders: true, // X-RateLimit-* headers
legacyHeaders: false,
store: new RedisStore({ sendCommand: (...args) => redisClient.sendCommand(args) }),
message: { error: 'Too many requests', retryAfter: 900 },
})
// Strict limit for auth endpoints
const authLimit = rateLimit({
windowMs: 15 * 60 * 1000,
max: 10,
skipSuccessfulRequests: true, // only count failed attempts
message: { error: 'Too many login attempts' },
})
app.use(globalLimit)
app.post('/auth/login', authLimit, loginHandler)
app.post('/auth/register', authLimit, registerHandler)
Before deploying this behind a load balancer or reverse proxy, check req.ip. By default Express reports the proxy’s address, so every client shares one bucket and the whole site gets rate limited together as soon as traffic picks up. Recent versions of express-rate-limit detect this and log a validation warning, but only if someone reads the logs. Set app.set('trust proxy', 1) (the number of proxy hops you actually have), not true. true trusts the left-most X-Forwarded-For entry, which the client controls.
Also notice that authLimit uses skipSuccessfulRequests. For a login endpoint, you care about failed attempts, not about a user who logs in successfully ten times. This does not weaken brute-force protection: successful logins are skipped, but they do not reset the failure count, so guessing still hits the limit. Newer versions of the library also accept limit as the name for max; both work.
Custom middleware with per-user limits
// Different limits per subscription tier
const TIER_LIMITS = {
free: { requests: 100, window: 3600 }, // 100/hour
pro: { requests: 1000, window: 3600 }, // 1000/hour
enterprise: { requests: 10000, window: 3600 }, // 10000/hour
}
async function rateLimitMiddleware(req, res, next) {
const user = req.user // set by auth middleware
const tier = user?.subscriptionTier ?? 'free'
const { requests, window } = TIER_LIMITS[tier]
const key = user ? `user:${user.id}` : `ip:${req.ip}`
const { allowed, count, retryAfter } = await slidingWindowLog(key, requests, window)
// Set standard headers
res.setHeader('X-RateLimit-Limit', requests)
res.setHeader('X-RateLimit-Remaining', Math.max(0, requests - count))
res.setHeader('X-RateLimit-Reset', Math.floor(Date.now() / 1000) + window)
if (!allowed) {
res.setHeader('Retry-After', retryAfter ?? window)
return res.status(429).json({
error: 'Rate limit exceeded',
limit: requests,
window: `${window}s`,
retryAfter: retryAfter ?? window,
})
}
next()
}
Two details are easy to miss. First, slidingWindowLog returns no retryAfter, so the handler falls back to the full window length. That is safe but pessimistic. A precise value would be the time until the oldest entry in the sorted set expires (ZRANGE key 0 0 WITHSCORES). Second, X-RateLimit-Reset is computed as “now + window”, which is only right for fixed windows. For sliding and token-bucket limiters, “reset” is fuzzy by nature. Pick a definition, document it, and keep it consistent, because client SDKs will be written against whatever you send.
The bigger question this middleware raises is what happens when Redis is down. The await throws, and unless you catch it, every request fails with a 500. Your rate limiter has become a single point of failure for the whole API. Most teams choose to fail open: catch the error, log it, emit a metric, and let the request through. A few minutes without limits is almost always better than a full outage. The exception is endpoints where the limit is a security control, such as login or password reset. Failing closed there is defensible.
Per-Endpoint Rate Limiting
// Different limits for different operations
const limits = {
// Expensive operations
'/api/generate-report': { requests: 5, window: 3600 }, // 5/hour
'/api/export': { requests: 10, window: 3600 }, // 10/hour
// API key endpoints
'/api/v1/': { requests: 10000, window: 3600 }, // 10k/hour
// Public endpoints
'/api/posts': { requests: 500, window: 3600 }, // 500/hour
}
function getLimit(path) {
for (const [pattern, limit] of Object.entries(limits)) {
if (path.startsWith(pattern)) return limit
}
return { requests: 1000, window: 3600 } // default
}
Prefix matching like this depends on the order of the object’s keys. /api/v1/ sits above /api/posts, which is harmless here, but if someone later adds /api/ above /api/export, the export limit silently stops applying. Sorting patterns by length (longest first), or matching on your router’s route names instead of raw paths, avoids that. Also include the limit’s name in the Redis key (user:42:export), not just the user. Otherwise every endpoint shares one counter, and a user who spent their export budget cannot read posts either.
Handling Rate Limit Responses (Client Side)
// JavaScript fetch client with automatic retry
async function fetchWithRetry(url, options = {}, maxRetries = 3) {
for (let attempt = 0; attempt < maxRetries; attempt++) {
const response = await fetch(url, options)
if (response.status === 429) {
const retryAfter = parseInt(response.headers.get('Retry-After') ?? '60')
console.warn(`Rate limited. Retrying in ${retryAfter}s`)
await new Promise(r => setTimeout(r, retryAfter * 1000))
continue
}
return response
}
throw new Error('Max retries exceeded')
}
// Axios interceptor
axios.interceptors.response.use(null, async (error) => {
if (error.response?.status === 429) {
const retryAfter = error.response.headers['retry-after'] ?? 60
await new Promise(r => setTimeout(r, retryAfter * 1000))
return axios.request(error.config)
}
return Promise.reject(error)
})
Both clients are fine for scripts but risky at scale. The Axios interceptor retries forever: nothing counts attempts, so a server that keeps returning 429 keeps the client looping. And both clients wait exactly Retry-After seconds. If a thousand clients were throttled at the same moment, they all come back at the same moment, get throttled again, and the load arrives in synchronized waves. Add a retry cap (for example, a counter on error.config) and random jitter to the wait: retryAfter * 1000 * (1 + Math.random() * 0.5).
Also note that Retry-After may legally be an HTTP date instead of a number of seconds. parseInt on a date string returns NaN, and setTimeout(r, NaN) fires immediately. That turns “back off” into “retry at once”. If you consume third-party APIs, handle both forms.
A lot of apparent “rate limit” incidents I have looked into turn out to be retry storms rather than too much real traffic. The pattern is always similar: a downstream service slows down, clients time out and retry, the retries add load, which causes more timeouts. Every client is behaving “correctly” according to its own retry policy. Setting server-side limits helps, but the lasting fix is on the client side: bounded retries with exponential backoff and jitter. That is why I now treat a good 429 contract (accurate Retry-After, documented headers) as part of the API design, not an afterthought.
Choosing the Limit Key and Where to Enforce It
IP-based vs user-based
function getRateLimitKey(req) {
// Authenticated users: rate limit by user ID (fair per account)
if (req.user?.id) return `user:${req.user.id}`
// API keys: rate limit by key
if (req.headers['x-api-key']) return `apikey:${req.headers['x-api-key']}`
// Unauthenticated: rate limit by IP
// Note: use X-Forwarded-For carefully behind proxies
const ip = req.headers['x-forwarded-for']?.split(',')[0].trim() ?? req.ip
return `ip:${ip}`
}
The IP branch above has a real hole: X-Forwarded-For is a client-supplied header. If your app is reachable directly, or your proxy appends to the header instead of replacing it, an attacker can send X-Forwarded-For: <random IP> on every request and get a fresh bucket each time. The first entry is the one the client controls. The trustworthy entry is the one added by your outermost proxy, counted from the right. Use Express’s trust proxy setting and read req.ip, rather than parsing the header yourself.
Two more notes on keys. Do not put the raw API key into the Redis key. Anyone with KEYS/SCAN access (or a Redis dump) then sees live credentials, so hash it first. And for IPv6, limit by prefix (typically a /64) rather than the full address. One home connection can easily rotate through millions of addresses in its range.
Bypass for trusted clients
const TRUSTED_IPS = new Set(['10.0.0.1', '10.0.0.2'])
function rateLimitMiddleware(req, res, next) {
if (TRUSTED_IPS.has(req.ip)) return next()
if (req.user?.isAdmin) return next()
// ... apply rate limiting
}
Cloudflare / Nginx rate limiting (infrastructure layer)
# nginx.conf — coarse rate limiting before app
limit_req_zone $binary_remote_addr zone=api:10m rate=100r/m;
location /api/ {
limit_req zone=api burst=20 nodelay;
limit_req_status 429;
proxy_pass http://app;
}
Nginx’s limit_req is a leaky bucket per worker-shared zone. rate=100r/m means one request every 600 ms, not “100 at any time within a minute”. Without burst, a client that sends two requests 100 ms apart gets the second one rejected, which surprises people the first time. burst=20 nodelay allows 20 requests over the steady rate to go through immediately, then enforces the rate. The limit is also per Nginx node. With three load-balanced Nginx servers, the effective limit is roughly 3×. That is fine for a coarse DDoS shield, and it is one reason precise per-user limits belong in the app with shared Redis state. See Nginx Reverse Proxy Configuration for the surrounding proxy setup.
Response Headers Reference
X-RateLimit-Limit: 1000 # requests allowed in window
X-RateLimit-Remaining: 847 # requests remaining
X-RateLimit-Reset: 1713312000 # Unix timestamp when limit resets
Retry-After: 3600 # seconds until client can retry (on 429)
Which Algorithm Fits Which Traffic
| Algorithm | Best for |
|---|---|
| Fixed window | Simple, high-traffic counters |
| Sliding window log | Accurate limits, lower traffic |
| Token bucket | APIs that allow bursts (recommended default) |
| Pattern | Recommendation |
|---|---|
| State store | Redis (shared across servers) |
| Rate by | User ID > API key > IP |
| Status code | 429 with Retry-After header |
| Infrastructure | Nginx/Cloudflare for DDoS, app for per-user |
| Different limits | Per endpoint, per tier, per operation type |
Implement rate limiting early — retrofitting it onto a production API is painful. Start with express-rate-limit + Redis for simplicity, then implement custom token bucket logic if you need tier-based limits or burst control.
Related Articles
- Redis Internals and Usage: The Event Loop, Encodings, RDB vs AOF, Replication and Cluster
- JWT Authentication Guide | Access Tokens· Refresh Tokens
- Node.js + Nginx Reverse Proxy Setup