Finding Memory Leaks: Heap Snapshots in Node.js, tracemalloc in Python, Heap Dumps in Java

Key takeaways

Memory leaks cause servers to grow until they crash. This guide gives you a systematic approach — using profilers and heap snapshots — to find and fix leaks in Node.js, Python, and Java.

Memory Leak vs. High Memory Usage

A memory leak is memory that is allocated but never released because references are retained unintentionally. Symptoms:

  • Memory usage increases monotonically over time
  • GC pauses become longer and more frequent
  • Service eventually crashes (OOM / exit code 137 in containers)

High memory usage that plateaus is not a leak — it may be a sizing issue but doesn’t require leak-hunting techniques.

The distinction is harder to see than it sounds, because garbage-collected runtimes rarely give memory back to the operating system promptly. V8, CPython’s allocator, and the JVM all keep freed memory reserved for future allocations, so RSS measured from outside often climbs during a traffic spike and then stays flat at the new level. That is a plateau, not a leak. The signal to look for is the floor after garbage collection: if the lowest point of heap usage after each GC cycle keeps rising under steady load, something is being retained. A sawtooth pattern whose valleys stay level is healthy no matter how high the peaks are.

The most common false alarm I have run into is a cache that is working as designed — it fills up over the first few hours, memory climbs steadily, and it looks exactly like a leak on a dashboard until it reaches its size limit and levels off. Watching for long enough (a full day of traffic, including the quiet hours) usually separates the two before any snapshot is needed.


General Approach

1. Confirm the leak — monitor memory over time
2. Isolate the scenario — which request/operation triggers growth?
3. Take heap snapshots before and after
4. Find objects growing between snapshots
5. Trace references back to root cause
6. Fix, deploy, confirm memory stabilizes

Node.js

Monitoring memory

// Log memory usage every 30 seconds
setInterval(() => {
  const m = process.memoryUsage();
  console.log({
    rss: `${Math.round(m.rss / 1024 / 1024)}MB`,       // total process memory
    heapUsed: `${Math.round(m.heapUsed / 1024 / 1024)}MB`,
    heapTotal: `${Math.round(m.heapTotal / 1024 / 1024)}MB`,
    external: `${Math.round(m.external / 1024 / 1024)}MB`,
  });
}, 30000);

Watch heapUsed — if it climbs continuously it confirms a heap leak. If rss climbs but heapUsed doesn’t, the leak may be in native code or external (Buffers).

Heap snapshots with Chrome DevTools

node --inspect app.js

Open Chrome → chrome://inspect → click “inspect” → Memory tab → Take heap snapshot.

Take a snapshot, trigger the suspected leak (run 100 requests, etc.), take another snapshot. In the second snapshot, select “Comparison” view to see what grew.

A more reliable variant is the three-snapshot technique: warm the app up first (so caches, JIT-compiled code, and lazy singletons are already allocated), take snapshot 1, run the scenario N times, take snapshot 2, run it N more times, take snapshot 3. Objects allocated between 1 and 2 that are still alive in 3 are your candidates; one-time initialization noise disappears. DevTools forces a garbage collection before each snapshot, so anything that remains is genuinely reachable.

In the snapshot, sort by Retained Size, not Shallow Size. Shallow size is the object’s own memory; retained size is everything that would be freed if that object were collected. A leak usually shows up as thousands of small objects (closures, Arrays, strings) whose retainers all lead back to one long-lived object — an EventEmitter’s listener array, a module-level Map, a Socket. Select a leaked object and read the Retainers panel bottom-up until you reach something your code owns; that is where the fix goes.

Heap snapshots programmatically

import v8 from 'v8';
import fs from 'fs';

// Take a snapshot and write to file (returns the file name)
const snapshotFile = v8.writeHeapSnapshot();
console.log('Snapshot written to:', snapshotFile);

// Or, without code changes (Node 12+): start with
//   node --heapsnapshot-signal=SIGUSR2 app.js
// and run `kill -USR2 <pid>` to write a .heapsnapshot
// (the old `heapdump` npm package did the same for older Node versions)

Taking a snapshot in production has real costs. writeHeapSnapshot is synchronous: the event loop stops for the whole duration, which is seconds for a heap of a few hundred megabytes, so health checks can fail and the load balancer may pull the instance. Serializing the heap also needs extra memory — the Node docs warn it can require roughly twice the heap size — which means a process that is already close to its container limit can be OOM-killed by the snapshot itself. Take snapshots from an instance that is out of rotation, or reproduce the leak in staging where you can raise the limit. For crash-time evidence, --heapsnapshot-near-heap-limit=1 writes a snapshot automatically when V8 approaches its heap limit.

Common Node.js leak patterns

1. Event listeners not removed

// LEAK: adds a listener on every request, never removes it
app.get('/data', (req, res) => {
  emitter.on('update', (data) => {  // grows unboundedly
    res.json(data);
  });
});

// FIX: use .once() or remove the listener
app.get('/data', (req, res) => {
  emitter.once('update', (data) => {
    res.json(data);
  });
});

Node gives an early warning for this pattern: after 10 listeners for the same event it prints MaxListenersExceededWarning: Possible EventEmitter memory leak detected. 11 update listeners added to [EventEmitter]. Do not silence it with setMaxListeners(0) without understanding why the count grows. Note that .once() only helps if the event eventually fires; if a client disconnects before update arrives, the listener (and the closed res it captures) stays attached. The robust fix removes the listener on req.on('close', ...) or uses events.once(emitter, 'update', { signal }) with an AbortController tied to the request lifetime.

2. Growing cache with no eviction

// LEAK: cache grows forever
const cache = new Map();
app.get('/user/:id', async (req, res) => {
  if (!cache.has(req.params.id)) {
    cache.set(req.params.id, await fetchUser(req.params.id));
  }
  res.json(cache.get(req.params.id));
});

// FIX: use LRU cache with size limit
import { LRUCache } from 'lru-cache';
const cache = new LRUCache({ max: 1000, ttl: 1000 * 60 * 5 });

The unbounded version is the most common leak in long-running services because it is invisible in development: with a few test users the map never gets big. In production, the key space is whatever users send — every distinct id, including invalid ones and those generated by scanners. A count limit (max) bounds the number of entries but not their size; if values vary widely, maxSize with a sizeCalculation function bounds the bytes. And a per-process cache multiplies by the number of processes or pods, which is often the point at which Redis becomes the better home for it.

3. Closures holding large objects

// LEAK: data (potentially large) is captured in the closure
function processLargeFile(filePath) {
  const data = fs.readFileSync(filePath);  // large buffer
  return setInterval(() => {
    console.log('File size:', data.length);  // data never released
  }, 1000);
}

// FIX: extract only what you need
function processLargeFile(filePath) {
  const size = fs.statSync(filePath).size;  // just the number
  return setInterval(() => {
    console.log('File size:', size);
  }, 1000);
}

4. Timers keeping references alive

// LEAK: interval keeps the object alive forever
class DataProcessor {
  constructor() {
    this.data = new Array(100000).fill('x');
    setInterval(() => this.process(), 1000);  // keeps `this` alive
  }
  process() { /* ... */ }
  destroy() {
    // No way to clean up — interval captures `this`
  }
}

// FIX: store interval reference and clear it
class DataProcessor {
  constructor() {
    this.data = new Array(100000).fill('x');
    this.interval = setInterval(() => this.process(), 1000);
  }
  process() { /* ... */ }
  destroy() {
    clearInterval(this.interval);
    this.data = null;
  }
}

Clinic.js (comprehensive profiling)

npm install -g clinic
clinic heapprofiler -- node app.js
# Run your load test, then Ctrl+C
# Opens a flamegraph showing allocations

The heap profiler answers a different question from snapshots: not “what is alive now” but “which code allocated the most”. That is useful when memory churn, not retention, is the problem (GC pauses from allocation-heavy code), and as a starting point when you do not know which request type is responsible. Combine it with a load generator such as autocannon so the profile reflects realistic traffic. It adds overhead, so run it against staging rather than production instances.


Python

Monitor memory with tracemalloc

import tracemalloc
import linecache

tracemalloc.start()

# ... run suspected leaky code ...

snapshot = tracemalloc.take_snapshot()
top_stats = snapshot.statistics('lineno')

print("Top 10 memory allocations:")
for stat in top_stats[:10]:
    print(stat)

A single tracemalloc snapshot shows where memory was allocated, which includes everything legitimate. For leak hunting, take two snapshots around the suspect workload and diff them with snapshot2.compare_to(snapshot1, 'lineno'); lines whose size keeps growing across repeated runs are the leak. tracemalloc.start(25) records 25 frames per allocation instead of one, which you need when the top line is inside a library (json/decoder.py) and you want to know which of your call sites caused it. Tracing slows allocation-heavy code noticeably, so enable it only while investigating. It also sees only allocations made through Python’s allocator — memory held by C extensions such as NumPy arrays allocated via malloc may not appear.

memory_profiler — line-by-line

pip install memory_profiler
from memory_profiler import profile

@profile
def my_function():
    data = [i for i in range(1000000)]
    result = sum(data)
    return result
python -m memory_profiler script.py

Output shows memory usage per line — spots exactly where allocations happen.

objgraph — find growing objects

pip install objgraph
import objgraph

# Show most common object types
objgraph.show_most_common_types(limit=20)

# Show what's growing between two points
objgraph.show_growth()

# Trace references to an object
objgraph.show_backrefs(objgraph.by_type('MyClass')[0], max_depth=5)

Common Python leak patterns

1. Mutable default argument

# LEAK: list is created once and reused across all calls
def add_item(item, items=[]):
    items.append(item)
    return items

# FIX
def add_item(item, items=None):
    if items is None:
        items = []
    items.append(item)
    return items

2. Circular references

# Cycles are not freed by reference counting; they wait for the cyclic GC.
# (Before Python 3.4, a __del__ method made the cycle uncollectable and it
# ended up in gc.garbage. PEP 442 fixed that in 3.4+.)
class Node:
    def __init__(self):
        self.other = None
    def __del__(self):
        pass

# FIX: use weakref for back-references so no cycle forms at all
import weakref

class Node:
    def __init__(self):
        self._other = None

    @property
    def other(self):
        return self._other() if self._other else None

    @other.setter
    def other(self, value):
        self._other = weakref.ref(value)

On modern Python, cycles are a latency and memory-peak issue rather than a true leak: the objects are freed, but only when the generational collector runs, not the moment the last reference disappears. That matters when the cycle holds something large (a parsed document, a frame with big locals) — memory stays high until the next collection. A genuine leak from cycles still happens when some object in the cycle is also referenced from a long-lived place, such as a module-level registry or a traceback stored in a logger. Storing sys.exc_info() or an exception object keeps its traceback alive, and the traceback keeps every frame’s local variables alive — a surprisingly common way to retain large objects in exception-handling code.

3. Django queryset caching

# LEAK in background task: qs evaluates and caches all rows
def process_all_users():
    users = User.objects.all()  # 1M rows loaded into memory
    for user in users:
        send_email(user)

# FIX: use iterator() to avoid caching
def process_all_users():
    for user in User.objects.all().iterator(chunk_size=1000):
        send_email(user)

Java / JVM

Heap dump

# Trigger heap dump on OOM automatically
java -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/tmp/heapdump.hprof -jar app.jar

# Trigger manually (find PID first)
jmap -dump:live,format=b,file=/tmp/heapdump.hprof <PID>
# Modern equivalent: jcmd <PID> GC.heap_dump /tmp/heapdump.hprof

The live option runs a full GC first so the dump contains only reachable objects — that is what you want for leak analysis, but the full GC plus the dump itself pause the application, often for seconds on large heaps. The .hprof file is roughly the size of the used heap, so make sure the target directory (often a small container filesystem) has room; a failed dump because /tmp filled up is a frustrating way to lose the evidence after waiting hours for the leak to reproduce. -XX:+HeapDumpOnOutOfMemoryError should be on by default in every JVM service: it costs nothing until the crash, and the dump taken at the moment of java.lang.OutOfMemoryError: Java heap space is usually the clearest picture of a leak you will get.

Analyze with Eclipse Memory Analyzer (MAT) or VisualVM:

  • “Leak Suspects Report” in MAT identifies the most likely leak
  • Look for objects with many retained instances

jstat — GC monitoring

# Monitor GC every 1 second
jstat -gcutil <PID> 1000

# Output columns: S0% S1% E% O% M% YGC YGCT FGC FGCT GCT
# O% = Old Gen usage — if this grows continuously, there's a leak

Read O together with FGC (full GC count). Old-generation usage rising between full GCs is normal; old-generation usage that stays high right after each full GC, while FGC increases faster and faster, is the leak signature. Near the end the JVM spends most of its time collecting and reclaims almost nothing, and you may see java.lang.OutOfMemoryError: GC overhead limit exceeded instead of the plain heap-space error.

Common Java leak patterns

1. Static collections

// LEAK: static map grows forever
public class Cache {
    private static final Map<String, Object> store = new HashMap<>();
    
    public static void put(String key, Object value) {
        store.put(key, value);  // never evicted
    }
}

// FIX: use Caffeine or Guava cache with eviction
import com.github.benmanes.caffeine.cache.Cache;
import com.github.benmanes.caffeine.cache.Caffeine;

Cache<String, Object> cache = Caffeine.newBuilder()
    .maximumSize(10_000)
    .expireAfterWrite(Duration.ofMinutes(5))
    .build();

2. Unclosed resources

// LEAK: connection never closed if exception is thrown
public void query() throws SQLException {
    Connection conn = dataSource.getConnection();
    Statement stmt = conn.createStatement();
    ResultSet rs = stmt.executeQuery("SELECT ...");
    // Exception here → conn never closed → connection leak
}

// FIX: try-with-resources
public void query() throws SQLException {
    try (Connection conn = dataSource.getConnection();
         Statement stmt = conn.createStatement();
         ResultSet rs = stmt.executeQuery("SELECT ...")) {
        // auto-closed even on exception
    }
}

3. ThreadLocal not cleaned up

// LEAK in thread pools: ThreadLocal value survives thread reuse
private static final ThreadLocal<LargeObject> holder = new ThreadLocal<>();

// FIX: always remove in finally
LargeObject obj = new LargeObject();
holder.set(obj);
try {
    processRequest();
} finally {
    holder.remove();  // critical in thread pool environments
}

Container / Kubernetes Context

# Watch memory usage of pods
kubectl top pods --sort-by=memory

# Check OOMKilled history (this event comes from node-level tooling and
# may not exist in every cluster; the pod status below is always there)
kubectl get events --field-selector reason=OOMKilling

# If container is OOMKilled, increase limit and check for leak
kubectl describe pod <pod-name>
# Look for: OOMKilled, Last State exit code 137

Set a memory limit in Kubernetes — this forces a crash (which you’ll notice) instead of unbounded growth:

resources:
  requests:
    memory: "256Mi"
  limits:
    memory: "512Mi"

Make sure the runtime knows about the limit. Node sizes its default old-space limit from available memory, which inside a container may or may not reflect the cgroup limit depending on the Node version; setting --max-old-space-size to roughly 75% of the container limit makes V8 collect aggressively and fail with a clear JavaScript heap out of memory error before the kernel kills the process silently with exit code 137. The JVM has been container-aware since Java 10 (and 8u191), and -XX:MaxRAMPercentage=75 is the usual way to size the heap relative to the limit. Leaving headroom matters because both runtimes use memory outside the heap — thread stacks, metaspace, buffers, native libraries.


Profiling Tools Summary

LanguageToolUse for
Node.jsChrome DevTools / --inspectHeap snapshots, allocation timeline
Node.jsClinic.jsAllocation profiling under load (staging)
Node.jsprocess.memoryUsage()Continuous monitoring
PythontracemallocAllocation by file/line
Pythonmemory_profilerLine-by-line memory usage
PythonobjgraphObject count growth
Javajmap + MATHeap dump analysis
Javajstat -gcutilGC health monitoring
AllPrometheus + GrafanaLong-term memory trend

Frequently Asked Questions (FAQ)

Q. My Node.js pod keeps getting OOMKilled, but heap snapshots look stable. What am I missing?

A. The Kubernetes memory limit applies to the whole container’s resident memory, not just the V8 heap that heap snapshots show. Buffers, native addons and other off-heap allocations count toward the limit too. Log process.memoryUsage() over time and compare rss and external against heapUsed: if rss climbs while heapUsed stays flat, the growth is outside the JS heap and you need to look at buffers and native code rather than JavaScript object references.