Node.js Performance: Clustering, Caching, and Profiling
Key takeaways
Most slow Node.js services are slow for one of three reasons: the database, the event loop being blocked, or work being repeated that could be cached. This guide covers how to tell which one you have, and the fixes (and their traps) for each.
Introduction
Performance work in Node.js goes wrong most often in the same way: someone adds clustering or a cache because it is on a checklist, the numbers barely move, and the real bottleneck — a query without an index, or a synchronous JSON parse of a 20 MB payload — is still there. So the order of this guide matters. Measure, find which of the usual suspects you actually have, fix that, and measure again.
The usual suspects, roughly in the order I find them in real services:
- The database — missing indexes, N+1 queries, fetching whole documents to use two fields.
- A blocked event loop — CPU-heavy work, synchronous APIs (
fs.readFileSync,crypto.pbkdf2Sync, bigJSON.parse) on the request path. While the loop is blocked, every request waits, not just the slow one. - Repeated work — the same expensive computation or remote call made for every request.
- Only then — using more cores, compression, connection tuning.
A useful first signal is event loop delay. If p99 latency is high but the event loop is idle, you are waiting on I/O (usually the database). If event loop delay is also high, something is hogging the thread.
// Built-in since Node 11: histogram of event loop delay
const { monitorEventLoopDelay } = require('perf_hooks');
const h = monitorEventLoopDelay({ resolution: 20 });
h.enable();
setInterval(() => {
// values are in nanoseconds
console.log(`loop delay p50=${(h.percentile(50) / 1e6).toFixed(1)}ms p99=${(h.percentile(99) / 1e6).toFixed(1)}ms`);
h.reset();
}, 10000);
Clustering
A Node.js process runs your JavaScript on a single thread. On an 8-core machine, one process uses roughly one core for JS, no matter how much load it gets. The cluster module forks worker processes that share a listening port, so the OS spreads connections across them.
// cluster.js
const cluster = require('cluster');
const http = require('http');
const os = require('os');
// availableParallelism() (Node 18.14+) respects CPU affinity; cpus().length does not
const numWorkers = os.availableParallelism ? os.availableParallelism() : os.cpus().length;
if (cluster.isPrimary) { // isMaster is the deprecated name
console.log(`Primary ${process.pid} running, forking ${numWorkers} workers`);
for (let i = 0; i < numWorkers; i++) cluster.fork();
cluster.on('exit', (worker, code, signal) => {
console.log(`Worker ${worker.process.pid} exited (${signal || code}), starting a new one`);
cluster.fork();
});
} else {
http.createServer((req, res) => {
res.writeHead(200);
res.end(`Handled by worker ${process.pid}\n`);
}).listen(3000, () => console.log(`Worker ${process.pid} started`));
}
Things clustering does not give you for free:
- Shared memory. Each worker has its own heap. An in-memory cache, rate limiter counter or WebSocket session map exists once per worker, so a user may hit a cold cache on worker 3 right after warming it on worker 1. Anything that must be shared across workers belongs in Redis or the database.
- A fix for a blocked loop. If one request blocks a worker for two seconds, clustering just means the other workers keep serving; that worker’s queued requests still wait.
- Blind auto-restart safety. The
exithandler above restarts forever. If a worker crashes on startup (bad config, port conflict), you get a tight fork loop that pegs the CPU. In production, add a backoff or let a process manager handle restarts.
One mistake I have made myself is running cluster mode inside a container that Kubernetes was also scaling. os.cpus() reported the node’s 32 cores, not the container’s 2-CPU limit, so each pod forked 32 workers that fought over 2 CPUs and used 16x the memory. In containers, one process per container and horizontal scaling is usually the cleaner model; if you do cluster, size it from the CPU limit, not the host.
PM2 cluster mode
PM2 does the same forking plus restarts, logs and zero-downtime reloads, which is why most bare-VM deployments use it instead of hand-written cluster code:
pm2 start app.js -i max # one process per available core
pm2 start app.js -i 4 # fixed count
pm2 reload app # restart workers one at a time, no dropped connections
pm2 reload is only zero-downtime if your app shuts down gracefully — stops accepting connections and finishes in-flight requests on SIGINT. Otherwise requests in flight on the old worker are cut off. See the PM2 guide for the graceful shutdown pattern.
Caching
Caching turns repeated work into a lookup, but every cache adds a new question: when is the cached value wrong? Before adding one, decide the TTL and what invalidates it. If you cannot answer that, you are trading a performance problem for a correctness bug.
In-memory cache (and the stampede problem)
// TTL cache that shares one in-flight promise per key
const cache = new Map();
async function getCached(key, fetchFn, ttlMs = 60000) {
const hit = cache.get(key);
if (hit && Date.now() < hit.expiresAt) return hit.promise;
const promise = fetchFn();
cache.set(key, { promise, expiresAt: Date.now() + ttlMs });
// Do not keep a failed fetch cached for the full TTL
promise.catch(() => cache.delete(key));
return promise;
}
app.get('/api/users', async (req, res) => {
const users = await getCached('users', () => User.find().lean(), 60000);
res.json({ users });
});
Storing the promise rather than the resolved value is deliberate. The naive version (check cache, await the database, then store the result) has a gap: if 500 requests arrive while the key is missing, all 500 miss and all 500 query the database. That is a cache stampede, and it tends to happen at the worst moment — right after a deploy empties the cache, or when a hot key expires at peak traffic. With the promise stored immediately, requests 2 through 500 await the same query as request 1.
The .catch removal matters too. Without it, one database timeout gets cached and every request for the next minute receives the same rejected promise.
Redis caching
Redis gives you a cache shared by all workers and all instances, and it survives an app restart. The cost is a network round trip (usually well under a millisecond on the same network) and serialization.
npm install redis
// cache.js — node-redis v4+
const { createClient } = require('redis');
// v4 takes a URL (or { socket: { host, port } }); the old { host, port } form is silently ignored
const client = createClient({ url: process.env.REDIS_URL || 'redis://localhost:6379' });
client.on('error', (err) => console.error('Redis error:', err));
async function connect() {
await client.connect();
}
async function setCache(key, value, ttlSeconds = 3600) {
// jitter spreads expiry so keys written together do not all expire together
const jitter = Math.floor(Math.random() * ttlSeconds * 0.1);
await client.set(key, JSON.stringify(value), { EX: ttlSeconds + jitter });
}
async function getCache(key) {
const value = await client.get(key);
return value ? JSON.parse(value) : null;
}
async function deleteCache(key) {
await client.del(key);
}
module.exports = { connect, setCache, getCache, deleteCache };
A detail that bit me: with node-redis v3, createClient({ host, port }) worked. After upgrading to v4, the same code connected to localhost:6379 regardless of what host said, because v4 expects url or socket.host. Locally everything passed; in staging, where Redis was on another host, every request failed with ECONNREFUSED. If you upgrade the client, check the constructor options, not just the command names.
const { getCache, setCache } = require('./cache');
app.get('/api/users/:id', async (req, res) => {
const cacheKey = `user:${req.params.id}`;
const cached = await getCache(cacheKey);
if (cached) return res.json({ user: cached, cached: true });
const user = await User.findById(req.params.id).lean();
if (!user) return res.status(404).json({ error: 'User not found' });
await setCache(cacheKey, user, 3600);
res.json({ user, cached: false });
});
Decide what happens when Redis is down. For a pure cache, the right behavior is usually to log and fall through to the database, not to return 500 — wrap getCache in a try/catch and treat errors as misses.
Cache invalidation
app.put('/api/users/:id', async (req, res) => {
const user = await User.findByIdAndUpdate(req.params.id, req.body, { new: true }).lean();
await deleteCache(`user:${req.params.id}`);
res.json({ user });
});
Delete-on-write is simple and correct for single-key data. It breaks down for derived data: if /api/users (the list) is cached too, updating one user leaves the list stale until its TTL expires. Either invalidate every key that includes the entity, or accept a short TTL on lists. More patterns (write-through, cache-aside, tag-based invalidation) are in Redis caching patterns for Node.js.
Database optimization
In most web services I have profiled, the time goes to waiting on the database, not running JavaScript. That is why this section comes before async tricks or clustering.
Indexes
// MongoDB via Mongoose
userSchema.index({ email: 1 }, { unique: true });
userSchema.index({ name: 1, age: -1 }); // compound: order matters
const indexes = await User.collection.getIndexes();
console.log(indexes);
A compound index on { name: 1, age: -1 } helps queries that filter on name, or on name and age, but not queries on age alone — the leftmost field has to be used. Verify with .explain('executionStats'): if you see COLLSCAN and totalDocsExamined far above nReturned, the index is not being used. For PostgreSQL the equivalent is EXPLAIN ANALYZE looking for Seq Scan on large tables.
N+1 queries
// Slow: 1 query for posts + 1 query per post
const posts = await Post.find();
for (const post of posts) {
post.author = await User.findById(post.author);
}
// Better: populate batches the author lookup into one $in query
const posts = await Post.find()
.select('title content author createdAt')
.populate('author', 'name email')
.lean()
.limit(20);
N+1 is easy to miss because it is fast in development: 20 posts × 1 ms on localhost is 20 ms. In production, with the database 2 ms away and 200 posts on the page, it becomes 400 ms of pure round trips. Enabling query logging (mongoose.set('debug', true) or your ORM’s equivalent) during a single request makes it obvious.
.lean() returns plain objects instead of Mongoose documents, skipping hydration, getters and change tracking. It is noticeably faster for read-only responses, but you lose save(), virtuals and defaults — do not use it when you intend to modify and save the result.
Connection pools
// Mongoose
mongoose.connect(process.env.MONGO_URL, { maxPoolSize: 10, minPoolSize: 2 });
// node-postgres
const { Pool } = require('pg');
const pool = new Pool({ max: 20, idleTimeoutMillis: 30000 });
Bigger is not better. A pool of 20 per process × 8 cluster workers × 5 instances is 800 connections, which exceeds PostgreSQL’s default max_connections of 100 by a wide margin. When the pool is the bottleneck, requests queue waiting for a connection; when the database is the bottleneck, a bigger pool just queues them inside the database instead. Do the multiplication before raising the number.
Async optimization
Run independent work in parallel
// Sequential: total time is T1 + T2 + T3
async function sequential() {
const users = await User.find();
const posts = await Post.find();
const comments = await Comment.find();
return { users, posts, comments };
}
// Parallel: total time is roughly max(T1, T2, T3)
async function parallel() {
const [users, posts, comments] = await Promise.all([
User.find(),
Post.find(),
Comment.find(),
]);
return { users, posts, comments };
}
Promise.all rejects as soon as one promise rejects, but the others keep running — you just stop waiting for them. If partial results are acceptable, use Promise.allSettled.
Limit concurrency
// p-limit v4+ is ESM-only; use import, or pin p-limit@3 for require()
import pLimit from 'p-limit';
async function processFiles(files) {
const limit = pLimit(5);
return Promise.all(files.map((file) => limit(() => processFile(file))));
}
Unbounded Promise.all over a large array is a common way to take down your own dependencies: 10,000 files means 10,000 simultaneous file handles (hello, EMFILE: too many open files) or 10,000 concurrent API calls that get you rate-limited. Pick a limit based on what the downstream resource can handle, not on what feels fast.
Memory management
Inspect memory
function logMemoryUsage() {
const mb = (n) => `${Math.round(n / 1024 / 1024)} MB`;
const m = process.memoryUsage();
console.log({ rss: mb(m.rss), heapTotal: mb(m.heapTotal), heapUsed: mb(m.heapUsed), external: mb(m.external) });
}
setInterval(logMemoryUsage, 60000);
How to read it: heapUsed going up and down in a sawtooth is garbage collection working normally. A leak looks like the lowest point of the sawtooth rising over hours. rss growing while heapUsed is flat points outside the JS heap — Buffers, native addons, or memory fragmentation. And if the process is killed with JavaScript heap out of memory in a container, check whether V8’s heap limit is sized for the container; recent Node versions derive it from the container memory limit, older ones did not, and --max-old-space-size sets it explicitly.
Bound your caches
// Bad: grows forever, one entry per distinct id ever requested
const cache = new Map();
// Better: LRU with a size cap and TTL (lru-cache v10+ API)
const { LRUCache } = require('lru-cache');
const lru = new LRUCache({ max: 500, ttl: 1000 * 60 * 60 });
app.get('/api/data/:id', async (req, res) => {
let data = lru.get(req.params.id);
if (!data) {
data = await fetchData(req.params.id);
lru.set(req.params.id, data);
}
res.json(data);
});
Older tutorials use new LRU({ maxAge }); since lru-cache v7 the option is ttl, and since v10 the export is the named LRUCache class. With the old option name the cache has no expiry at all, only the size cap.
An unbounded Map keyed by user input is the most common leak I see in code review, because it looks like a harmless cache and passes every test. It only shows up after days in production, as memory that climbs until the process restarts.
Stream large payloads
const fs = require('fs');
const { pipeline } = require('stream/promises');
// Bad: whole file in memory per request
app.get('/download', async (req, res) => {
res.send(await fs.promises.readFile('large-file.pdf'));
});
// Good: stream, with error handling and cleanup
app.get('/download', async (req, res, next) => {
try {
res.type('application/pdf');
await pipeline(fs.createReadStream('large-file.pdf'), res);
} catch (err) {
if (!res.headersSent) next(err);
}
});
Prefer pipeline over .pipe(). With plain .pipe(), an error on the read stream (file missing) is not forwarded, and if the client disconnects mid-download the read stream is not destroyed — file descriptors leak slowly under real traffic where people cancel downloads.
Profiling
Guessing at hot spots is almost always wrong. Profile under realistic load (run a load test in another terminal while profiling), because a function that is fast in isolation can dominate under concurrency.
Built-in CPU profiler
# Writes a .cpuprofile you can open in Chrome DevTools (Performance tab) or VS Code
node --cpu-prof app.js
# Older V8 tick profiler, text output
node --prof app.js
node --prof-process isolate-0x*.log > processed.txt
--cpu-prof writes the profile when the process exits cleanly, so stop the server with Ctrl+C rather than killing it with SIGKILL. In the flame chart, look for wide bars that are your code or a library you call; a wide bar under (garbage collector) means allocation pressure — too many short-lived objects — rather than slow logic.
Chrome DevTools
node --inspect app.js # attach any time
node --inspect-brk app.js # pause before the first line
Open chrome://inspect, attach, and use the Performance tab for CPU and the Memory tab for heap snapshots. For leaks: take a snapshot, run load for a few minutes, take another, and use the Comparison view to see which constructors grew. Do not leave --inspect bound on a public interface in production — it allows arbitrary code execution.
clinic.js
npm install -g clinic
clinic doctor -- node app.js # first pass: is it I/O, event loop, GC, or memory?
clinic flame -- node app.js # CPU flame graph
clinic bubbleprof -- node app.js # where async time goes
clinic heapprofiler -- node app.js # allocations
clinic doctor is the useful starting point because it tells you which kind of problem you have before you pick a more specific tool. Note that clinic.js has seen little maintenance recently; if it fails on a new Node version, --cpu-prof and DevTools cover the same ground.
Benchmarking
npm install -g autocannon
autocannon -c 50 -d 20 http://localhost:3000/api/users
const autocannon = require('autocannon');
async function runBenchmark() {
const result = await autocannon({ url: 'http://localhost:3000', connections: 50, duration: 20 });
console.log('Req/sec (avg):', result.requests.average);
console.log('Latency p99 (ms):', result.latency.p99);
console.log('Non-2xx responses:', result.non2xx);
}
runBenchmark();
Look at p99 latency and error counts, not just requests per second. A change that raises throughput by 20% but doubles p99 is usually a regression for real users. Also check non2xx: a server returning fast 500s benchmarks very well.
Apache Bench (ab) still works, but it defaults to HTTP/1.0 without keep-alive, so each request opens a new TCP connection and you end up measuring connection setup. Use ab -k if you use it at all.
Benchmark on hardware and data sizes close to production. Running the load generator on the same machine as the server steals CPU from the thing you are measuring.
Practical examples
Example 1: response caching middleware
const express = require('express');
const { createClient } = require('redis');
const app = express();
const client = createClient({ url: process.env.REDIS_URL });
function cacheMiddleware(ttlSeconds = 60) {
return async (req, res, next) => {
if (req.method !== 'GET') return next();
const key = `cache:${req.originalUrl}`;
try {
const cached = await client.get(key);
if (cached) return res.json(JSON.parse(cached));
} catch (err) {
return next(); // Redis down: serve uncached
}
const originalJson = res.json.bind(res);
res.json = (data) => {
// only cache successful responses
if (res.statusCode >= 200 && res.statusCode < 300) {
client.set(key, JSON.stringify(data), { EX: ttlSeconds }).catch(() => {});
}
return originalJson(data);
};
next();
};
}
app.get('/api/users', cacheMiddleware(60), async (req, res) => {
res.json({ users: await User.find().lean() });
});
(async () => {
await client.connect(); // top-level await is not available in CommonJS
app.listen(3000);
})();
Two things to watch with URL-keyed caching. First, the status check: without it, one transient 500 gets cached and served for the full TTL. Second, the key must include everything that changes the response. If /api/me depends on the logged-in user and the key is only the URL, user B receives user A’s cached profile. Never put per-user responses behind a URL-only cache.
Example 2: image optimization with sharp
const sharp = require('sharp');
const multer = require('multer');
const upload = multer({ dest: 'uploads/', limits: { fileSize: 10 * 1024 * 1024 } });
app.post('/upload', upload.single('image'), async (req, res, next) => {
try {
const outputPath = `optimized/${req.file.filename}.jpg`;
await sharp(req.file.path)
.resize(800, 600, { fit: 'inside', withoutEnlargement: true })
.jpeg({ quality: 80, mozjpeg: true })
.toFile(outputPath);
res.json({ path: outputPath });
} catch (err) {
next(err);
}
});
sharp does its work on libuv’s thread pool, not the main thread, so it does not block the event loop. But that pool has only 4 threads by default (UV_THREADPOOL_SIZE), shared with fs, dns.lookup, crypto.pbkdf2 and zlib. A burst of image uploads can make unrelated file reads and DNS lookups wait in line. If uploads are frequent, move image processing to a background job queue.
Compression
npm install compression
const compression = require('compression');
app.use(compression({
level: 6,
threshold: 1024, // do not bother compressing tiny responses
filter: (req, res) => {
if (req.headers['x-no-compression']) return false;
return compression.filter(req, res);
},
}));
Text formats like HTML and JSON typically compress very well; already-compressed formats (JPEG, PNG, video, zip) do not and should not be recompressed. The trade-off is CPU: gzip runs in your Node process. If you have nginx, a CDN or a load balancer in front, let it compress and keep that work off the event loop — it is usually faster there and can serve Brotli too.
Common pitfalls
Blocking the event loop
// Bad: CPU-heavy work on the main thread — every other request waits
app.get('/heavy', (req, res) => {
let sum = 0;
for (let i = 0; i < 1e9; i++) sum += i;
res.json({ sum });
});
// Better: a worker pool (piscina), created once at startup
const Piscina = require('piscina');
const pool = new Piscina({ filename: require.resolve('./heavy-task.js') });
app.get('/heavy', async (req, res, next) => {
try {
res.json({ sum: await pool.run({ n: 1e9 }) });
} catch (err) {
next(err);
}
});
// heavy-task.js
module.exports = ({ n }) => {
let sum = 0;
for (let i = 0; i < n; i++) sum += i;
return sum;
};
Spawning new Worker() per request, as many examples do, works in a demo but costs tens of milliseconds and a fresh V8 isolate each time, and under load you get hundreds of threads. A pool created once and reused is the production version.
The less obvious blockers are the ones that do not look like loops: JSON.parse/JSON.stringify on multi-megabyte payloads, synchronous crypto (bcrypt.hashSync), regular expressions with catastrophic backtracking on user input, and *Sync fs calls inside request handlers. The event loop delay histogram from the introduction catches all of them.
Listener leaks
const EventEmitter = require('events');
const emitter = new EventEmitter();
// Bad: a new listener on every request, never removed
app.get('/api/data', (req, res) => {
emitter.on('data', (data) => res.json(data));
});
// Better: once() removes itself; also clean up if the client leaves first
app.get('/api/data', (req, res) => {
const handler = (data) => res.json(data);
emitter.once('data', handler);
req.on('close', () => emitter.off('data', handler));
});
Node warns with MaxListenersExceededWarning: Possible EventEmitter memory leak detected when more than 10 listeners are attached to one event. Treat that warning as a bug report, not noise to silence with setMaxListeners(0).
Keep-alive behind a load balancer
const server = app.listen(3000);
// Must be longer than the load balancer's idle timeout (AWS ALB default: 60s)
server.keepAliveTimeout = 65000;
server.headersTimeout = 66000; // must be greater than keepAliveTimeout
This one looks like a tuning detail, but the symptom is intermittent 502 errors. Node’s default keepAliveTimeout is 5 seconds; ALB keeps idle connections for 60. After 5 idle seconds Node closes the socket, and if the balancer sends a request down that connection at the same moment, the request fails with a 502 that never appears in your application logs. It shows up as a low, steady error rate that gets worse at low traffic, which makes it confusing to chase. Setting the Node timeout above the balancer’s fixes it.
Timing slow requests
app.use((req, res, next) => {
const start = process.hrtime.bigint();
res.on('finish', () => {
const ms = Number(process.hrtime.bigint() - start) / 1e6;
if (ms > 1000) console.warn(`Slow: ${req.method} ${req.originalUrl} ${ms.toFixed(0)}ms`);
});
next();
});
process.hrtime.bigint() is monotonic, so it is not affected by system clock adjustments the way Date.now() is. Logging only slow requests keeps the output small enough that people actually read it.
Symptom, likely cause, first thing to try
| Symptom | Likely cause | First thing to try |
|---|---|---|
| High latency, low CPU, low loop delay | Waiting on the database or a remote API | Query logging, explain(), indexes, caching |
| High loop delay, one core pegged | Synchronous/CPU work on the main thread | CPU profile, move work to a worker pool |
| Latency spikes after deploys or TTL expiry | Cache stampede | Share in-flight promises, TTL jitter |
| Memory floor rising over hours | Unbounded cache, listener leak, retained closures | Heap snapshot comparison |
| Random 502s behind a load balancer | keep-alive timeout mismatch | keepAliveTimeout above the LB idle timeout |
| One core busy, others idle, loop healthy | Single process | Cluster/PM2, or more container replicas |
The order of priority that has held up for me: fix the database first, keep the event loop free, cache with an invalidation plan, and only then add processes and tune the network. Measure after each change, one change at a time, so you know which one helped.