C++ Memory Leaks: Causes, Detection, and Prevention

Key takeaways

Heap leaks, tools (Valgrind memcheck, LeakSanitizer, Massif, Heaptrack, VS diagnostics), smart pointers, STL pitfalls, false positives, suppressions, CI, and production monitoring.

Related: leak detection walkthrough · Valgrind guide · smart pointers · circular references.

Introduction: when memory only grows

A memory leak in C++ is more than a single missed delete. It is a class of lifetime bugs in which heap memory becomes unreachable through any pointer the program could still use to delete or delete[] it. The practical consequence is a steady rise in resident set size (RSS) in long-running processes, or allocation failure once the process hits a limit or the system runs short of memory.

This post covers heap leaks, reference-counting cycles, and the tools used to find them: Valgrind, AddressSanitizer with LeakSanitizer, heap profilers, and the Visual Studio diagnostics. It does not cover GPU memory or custom pools that never go through new.


What are memory leaks?

At the language level, dynamic storage comes from new and delete (or new[] / delete[]), usually hidden inside smart pointers and containers. If a heap block can no longer be named through any live pointer, reference, or smart pointer that will run its destructor, it is leaked: no valid operation in the program can free it, and it stays allocated until the process ends and the operating system reclaims the whole address space.

At the operating system level, a leak increases commit or RSS depending on the allocator. The C++ runtime (often glibc malloc on Linux, or msvc CRT on Windows) requests pages from the OS as needed. Heap exhaustion occurs when a new allocation request cannot be satisfied: std::bad_alloc for new, or nullptr from new (nothrow) depending on the form. In long-running services, slow leaks produce gradual pressure; in batch jobs, a leak might be “benign” if the process exits before pressure matters—but it still masks ownership bugs and breaks tests that run LeakSanitizer or Valgrind, which expect no unfreed direct leaks at normal exit (depending on your policy).

Lost memory is a useful informal phrase: the program lost the last way to name that allocation. A dangling pointer to freed memory is a different bug (use-after-free). A leak is the opposite situation: the memory is still allocated, but no correct pointer to it remains.

Heap exhaustion in C++ is not only “bad_alloc immediately.” Many services crawl toward limits: the kernel cgroup memory cap, the per-process ulimit on address space, or swapping to death on mis-sized VMs. A leak can also interact with fragmentation so that a new of size N fails even though the summary of free lists suggests enough bytes in aggregate—malloc and operator new are not “sum of free = can allocate max block.”

Process memory: RSS versus heap size

RSS (resident set size) is OS-level: pages the process has in RAM (simplified; exact definitions vary by ps flags and OS). The C++ heap is a sub-account inside that world: glibc may not return freed pages to the OS for a long time, so after you fix a leak, RSS can stall high until the allocator returns arenas or the process restarts. In SRE postmortems, say “leak fixed in code, RSS normalized after deploy; allocator retained part of the heap until restart” when that pattern appears.

Virtual size can grow even when RSS is flat if you mmap or reserve large arenas; not every “big virtual” line item is a C++ new leak. Cross-check with pmap / /proc/self/smaps on Linux or VMMap-style tools on Windows when sanitizers and leak tools disagree with ops graphs.

Why leaks hurt production

  • OOM killer (Linux) or commit failures (Windows) under sustained growth.
  • Noisy monitoring: “memory is high” without a single obvious stack.
  • False confidence when development machines never run the process long enough, or only short runs are profiled.

A typical failure mode is a slow, weeks-long climb in RSS rather than a spike. Short CI test suites pass under LeakSanitizer because the leaking path, for example an error return on a long-lived connection, is never exercised long enough or at all. The process then reaches its cgroup memory limit in production and gets killed. Two lessons follow: a clean exit in CI does not prove a process is safe to run for a month, and after a leak is fixed, RSS may stay high for a while because the allocator keeps freed pages until it decides to return them.


Common causes

Each subsection below shows a minimal broken pattern, then a typical fix, usually RAII. In real code these patterns often appear together, for example an exception path inside a container of raw pointers touched by a background thread.

Forgetting delete or delete[]

The most direct leak: allocate and return without a paired free.

void leak_basic() {
    int* p = new int(42);
    (void)p;
    // missing: delete p;
}

delete[] must pair with new[]:

void leak_mismatch() {
    int* a = new int[10];
    delete a; // undefined behavior; also often leaks depending on platform
}

Fix: one dynamic object → delete once; array → delete[] once, or use std::vector, std::unique_ptr<T[]>, or std::array when size is known at compile time.

Exception before delete (or early return)

If control leaves a scope that allocated with raw new and no RAII, every path must free. Exceptions and early returns are easy to miss.

void maybe_throw(bool fail);

void leak_on_exception() {
    int* p = new int(1);
    maybe_throw(true);     // if this throws, delete never runs
    delete p;
}

Fix: std::unique_ptr<int> p{new int(1)}; (or auto p = std::make_unique<int>(1);).

The same story applies to multiple return statements without a cleanup section.

Lost pointers (reassigning the only copy)

The program still “has a pointer” in the variable name, but you reassign it before delete, orphaning the old block if it was the only handle.

void lost_pointer() {
    int* p = new int(1);
    p = new int(2); // first allocation leaked
    delete p;
}

Fix: again, a single owner type (std::unique_ptr) or a container that disposes the old value when the slot is overwritten, depending on design.

Circular references with std::shared_ptr

shared_ptr is not a silver bullet. If object A holds a shared_ptr to B and B holds a shared_ptr back to A, the reference count never drops to zero when the external world releases its shared_ptrs, because the cycle keeps counts ≥ 1. That is a logical leak of the whole subgraph.

A minimal pattern:

#include <memory>

struct B;

struct A {
    std::shared_ptr<B> b;
};

struct B {
    std::shared_ptr<A> a; // cycle: A -> B -> A
};

Fix: at least one direction should be std::weak_ptr, or you redesign ownership to a tree/DAG, or you use a parent context object with clear lifetime. See Circular references / weak_ptr.

Assignment without delete (raw pointer members)

A copy assignment or reassignment for a class that owns raw new memory must delete[] the old buffer before replacing the pointer, or your rule of three/five/zero is violated.

Broken sketch:

class BadBuf {
    int* p_;
    size_t n_;
 public:
    BadBuf() : p_(nullptr), n_(0) {}
    void reset(size_t n) {
        p_ = new int[n]{};  // old p_ leaked
        n_ = n;
    }
    // ...
};

Fixes: std::vector<int> data_; (preferred), or follow rule of zero with a single owning member, or std::unique_ptr<int[]> with std::exchange in reset to delete[] the previous buffer explicitly while keeping exception safety in mind (again, vector is simpler).


Detection tools

Flags change between Clang, GCC, and MSVC releases, so check the documentation of the version you use. The mental model stays the same: Valgrind instruments at runtime, ASan+LSan inject checks at compile time, and heap profilers earn their keep when “something is growing” but the fault is indirect or the growth is not one obvious new.

Valgrind memcheck (--leak-check=full)

Valgrind on Linux (typical) runs your program in a synthetic CPU; memcheck tracks every heap block and reports definitely / indirectly / possibly lost and still reachable at process exit, depending on your summary settings.

Typical incantation:

valgrind --leak-check=full --show-leak-kinds=all \
    --track-origins=yes --verbose ./your_program arg1
  • --leak-check=full: more detailed per-leak reports, stack traces for allocation site.
  • --show-leak-kinds=all: do not only show “definitely lost”; helps you triage.
  • --track-origins=yes: helps with uninit use in some cases; not always needed for pure leak triage, adds cost.
  • For children: you may need --trace-children=yes if your test fork/execs.

Interpreting output: you see definitely lost (no pointer in any register/stack/global, according to the tool’s model), indirect (a root pointer lost but blocks hanging off it are “indirectly” lost from that root), and still reachable (pointers still exist at exit—often one-time init caches, sometimes intentional). Valgrind is slow (order of magnitude), but needs no special build flags—handy for release binaries (with debug symbols for stacks: -g).

AddressSanitizer and LeakSanitizer (ASan/LSan)

AddressSanitizer (-fsanitize=address on Clang/GCC) finds out-of-bounds accesses, use-after-free, double free, and more. On Linux, LeakSanitizer is bundled with ASan and enabled by default: at program exit, it scans memory for pointers to live heap blocks and reports the unreferenced ones with their allocation stacks. Enable it with a sanitizer build:

# Typical Clang/GCC debug build
c++ -g -O1 -fsanitize=address -fno-omit-frame-pointer -std=c++17 \
    main.cpp -o a.out
./a.out

Notes:

  • LSan is very fast compared to Valgrind, but you must rebuild with the sanitizer. Line numbers in leaks are excellent if -g and -fno-omit-frame-pointer are on.
  • ThreadSanitizer (TSan) and ASan are not always combined the same way on every version; for targeted leak runs, some teams use a dedicated ASan+LSan build in CI, then TSan in another job, rather than a single “everything at once” binary, depending on compiler support and policies.
  • MSVC supports AddressSanitizer (/fsanitize=address) for out-of-bounds and use-after-free errors, but it does not include LeakSanitizer, so it does not report leaks. On Windows, leak reports come from the CRT debug heap or Visual Studio snapshots instead.
  • On macOS, leak detection in ASan depends on the toolchain and platform; Apple’s Clang does not enable it by default, so Instruments’ Leaks template is the usual tool there.

Runtime flags (varies, often via ASAN_OPTIONS):

  • ASAN_OPTIONS=detect_leaks=1 turns leak detection on where it is supported but not enabled by default.
  • ASAN_OPTIONS=exitcode=0 to avoid failing the test binary when you are only gathering reports (rare; usually you want failure in CI when leaks are present).

For leak reports, read the “direct leak” and “indirect” sections similarly to Valgrind: the top frames of the allocation stack are your first place to add ownership or fix lifetime.

Heap profilers: Massif, heaptrack, and when to use them

Massif (Valgrind tool) samples heap + stack usage over time. It is ideal when the question is how much and where growth happens, not only a single “forgot to delete at exit.”

valgrind --tool=massif ./your_long_running_scenario
ms_print massif.out.12345

You get time-series and peaks; then you look at the heavy allocation stacks.

heaptrack (Linux) is a fast native heap profiler that can attach to processes (depending on setup) and produces chronological views of allocations and leaked or freed blocks with stack traces, often with a GUI (heaptrack_gui). It is a strong middle ground: faster than memcheck for some workloads, and very informative for leaks in loops and cumulative waste.

When to prefer profilers over memcheck/LSan alone: growth over hours in a service, or fragmentation-like symptoms where a simple new is not the whole story, or when you need attribution to hot paths in performance-sensitive code.

A heap-profiler workflow (heaptrack on Linux)

  1. Build with debug symbols (-g); you do not need asan to record (unless you are cross-checking).
  2. Run heaptrack ./service -- --config=stress.ini (arguments after -- to your app), or use heaptrack -p PID to attach in controlled environments.
  3. Open the .zst result in heaptrack_gui: sort by leaked or by trend in “memory over time.”
  4. For each top stack, ask: is this a leak (no path to free until exit), a per-request cache that is unbounded, or a leak in third-party code?
  5. Fix, then re-run the same workload and diff the peak RSS in htop as a sanity check; if the slope is flat but RSS is still high, read the RSS versus heap size section again.

Massif is similar in spirit, but the massif.out.* + ms_print text report is an excellent archival artifact in tickets when you need repro without a GUI on a headless runner.

Instruments (macOS): the Leaks and Allocations templates give time-based and backtrace views comparable to heaptrack for local GUI and iOS-adjacent C++; export or screenshot the heaviest call trees for your bug system.

Visual Studio diagnostic tools (Windows)

On Windows with MSVC, the native tools are the Debug CRT heap checks and leak report, and the memory usage snapshots in Visual Studio’s diagnostic tools.

Practical points:

  • Run under the debugger with the default Debug heap, which can fill freed blocks to catch some use-after-frees, and provides leak reporting hooks when configured.
  • The “Memory usage” view (tooling name can vary with VS version) allows heap snapshots and diffs to see what grew between two moments while exercising a feature.
  • AddressSanitizer for MSVC catches memory corruption bugs but, as noted above, not leaks.

CRT debug heap (classic snippet pattern—include only in debug, guard with your macros as appropriate):

#ifdef _DEBUG
#include <crtdbg.h>
// At startup (once):
// _CrtSetDbgFlag(_CRTDBG_ALLOC_MEM_DF | _CRTDBG_LEAK_CHECK_DF);
#endif

This asks for a leak report at process exit in debug builds. The report lists unfreed blocks by allocation number. _CRTDBG_MAP_ALLOC adds file and line information for malloc-family calls; for new, the usual trick is a debug-only macro that maps new to new(_NORMAL_BLOCK, __FILE__, __LINE__).

Dr. Memory (Windows/Linux in some modes) is another lightweight native alternative when you need realloc-style and uninitialized read help alongside leak reports; compare against your version matrix before standardizing in CI. Many teams use it when WSL2 is not an option and ASan is not yet in their MSVC path.

Valgrind: machine-readable output and suppressions

For Jenkins, GitLab, or in-house dashboards, emit XML so jobs can fail on definite only:

valgrind --xml=yes --xml-file=valgrind.xml \
    --leak-check=full --show-leak-kinds=definite \
    ./test_runner

--error-exitcode=1 makes the process exit non-zero on any memcheck error (including definite leaks, depending on flags); combine with your CI parser’s policy so still reachable does not fail a gate if you choose to allow it.

Generate a first-cut suppression interactively: run the leak once with --gen-suppressions=all and paste the block the tool prints into a .supp file, then trim wildcards to avoid over-matching. Review every suppression in code review; treat it like a // NOLINT in spirit.

LeakSanitizer: more ASAN_OPTIONS and LSAN_OPTIONS

Typical variables (read your exact clang docs—names evolve):

  • report_objects=1 (when supported): include per-object detail in the report.
  • use_unaligned=0 / allocator quirks: only change when your toolchain documents it.
  • suppressions=path for LSan (parallel to Valgrind’s file) to allow-list allocations in third-party code.

If you use LD_PRELOAD, custom allocators, or substitute malloc, LSan can false-negative or false-positive depending on how the allocator registers regions—stick to the default CRT/malloc in sanitizer test configs unless you are debugging allocator integration on purpose.

AddressSanitizer without LeakSanitizer (when split)

On some embedded or weird build graphs, you might compile ASan without the leak scan at exit to save a few seconds in tight inner loops, then use a dedicated LSan target for exit-time heap accounting. In practice, most desktop/server teams keep them together; splitting is an advanced optimization with paper trail in CMakePresets or Bazel constraint_values.


Smart pointers, RAII, and factories

RAII (Resource Acquisition Is Initialization) ties cleanup to scope: an owner’s destructor runs on scope exit, including exception paths, and releases heap memory, file handles, locks, and other resources. For heap objects, smart pointers and containers are the usual owners.

std::unique_ptr

Exclusive ownership, move-only, zero or negligible overhead. Prefer std::make_unique<T>(args...) (C++14) to avoid a transient raw pointer and to make exception safety in compound expressions less subtle.

#include <memory>

void good() {
    auto p = std::make_unique<int>(42);
    // p destroyed here; memory freed
}

For arrays, either std::vector, or std::make_unique<Widget[]>(n) (C++14 and up for std::make_unique array forms—verify your standard library implementation), depending on the case.

std::shared_ptr and std::weak_ptr

Shared ownership: the last shared_ptr to a control block destroys the managed object. Cycles are the classic shared_ptr footgun—see Common causes: circular references.

weak_ptr: non-owning observer; you lock() to a temporary shared_ptr if the object still exists. It is the right tool when a child needs to “see” a parent that might go away, or to break back-edges in graphs.

std::make_shared

std::make_shared<T>(args...) (C++11) combines allocation of the object and the control block when possible, improving locality and exception safety compared to std::shared_ptr<T>(new T(args...)), though custom deleters and weak_ptr lifetime subtleties may affect when make_shared is appropriate—verify against your C++ version and the object size if you are extremely allocation-sensitive.

Rule of thumb for leak prevention: no bare new in application code except at low-level library boundaries, and there immediately hand off to a smart pointer or standard container. Enforce with code review and linters where possible.

Custom deleters and C APIs

std::unique_ptr carries a deleter type; use it to pair malloc/free, new[]/delete[], or XFree-style functions without leaking the rule across the codebase.

#include <cstdlib>
#include <memory>

struct MallocDeleter {
    void operator()(void* p) const noexcept { std::free(p); }
};

using CBufferPtr = std::unique_ptr<void, MallocDeleter>;

CBufferPtr take_ownership(void* p) { return CBufferPtr(p); }

shared_ptr also accepts a deleter in its constructor, but the deleter is type-erased and has a small runtime cost. Prefer unique_ptr for sole ownership of non-shared C resources.

std::shared_ptr control blocks: make_shared trade-offs (brief)

std::make_shared can place the T and the control block in a single heap allocation. When the last shared_ptr is destroyed, T runs its destructor, but the block (including sizeof(T)-sized space for a destroyed object) can remain reserved until the last weak_ptr to that control block goes away. That is not a leak in the LSan/Valgrind sense, but in embedded or footprint-critical code it can be surprising. For server projects, the dominant issues remain retained graphs and cycles—not this—unless you profile and see weak_ptr-heavy designs holding large arenas alive.

Arrays: prefer std::vector to raw new[]

A raw new[] of non-trivial Widget is easy to get wrong: any throw between allocation and delete[] (or a missing delete[] on one path) leaks the whole array. std::vector<Widget>, with constructors and destructors run by the container, pairs naturally with RAII and with sanitizers. If you preallocate, use reserve plus push_back/emplace_back, or default-construct N elements only if that is what you need; shrink_to_fit is for rare peak-to-idle footprint tuning, not a primary leak tool.


Container management

STL containers are leak-safety for value types: std::vector<Widget>, std::map<K,V> (values), etc., destroy their elements in destructors. Leaks show up when you store pointers (especially raw pointers) and never delete what they point to.

Bad:

#include <vector>

void leak_in_vector() {
    std::vector<Widget*> v;
    v.push_back(new Widget{});
} // Leaks: vector destroys pointer values, not pointees

Good (often): std::vector<std::unique_ptr<Widget>> or std::vector<Widget> if copy semantics are defined and affordable.

Polymorphic base pointers: it is the owner’s responsibility to delete through a virtual destructor; prefer unique_ptr with custom deleters or shared_ptr with clear ownership story.

POD arrays in containers: std::array and std::vector of trivial types do not leak, but if you resize in ways that keep orphaned indirection, that is a design issue—usually use contiguous values, not new per element, unless a profiling reason forces it.


Custom allocators and tracking

For teaching, tracing, and controlled experiments, a stateful Allocator type passed to std::vector (or a pool type) can log every allocate/deallocate call. A leak in custom logic (double allocate, return wrong pointer) then becomes a mismatch in logs at shutdown.

A minimal C++-style counter pattern (pseudocode-level; align with your standard library’s Allocator requirements if you go production-grade):

#include <cstddef>
#include <memory>

struct tracking_allocator_stats {
    std::size_t live = 0;
};

template <class T>
struct tracking_allocator {
    using value_type = T;
    explicit tracking_allocator(tracking_allocator_stats& s) noexcept : s(&s) {}
    // Converting constructor: node-based containers rebind the allocator to their node type
    template <class U>
    tracking_allocator(const tracking_allocator<U>& o) noexcept : s(o.s) {}

    T* allocate(std::size_t n) {
        s->live += n;
        return std::allocator<T>{}.allocate(n);
    }
    void deallocate(T* p, std::size_t n) noexcept {
        s->live -= n;
        std::allocator<T>{}.deallocate(p, n);
    }
    template <class U>
    bool operator==(const tracking_allocator<U>& o) const { return s == o.s; }
    template <class U>
    bool operator!=(const tracking_allocator<U>& o) const { return s != o.s; }  // needed before C++20
    tracking_allocator_stats* s;
};

Use cases: test suites asserting stats.live == 0 at the end of a scope, or correlating a feature toggle with unexpected heap growth. This only counts allocations made through this allocator; it does not replace LSan or Valgrind.


Leak patterns

Patterns that recur in code reviews and incident reports:

  • “One per request” allocation without a matching free when the request fails mid-way (half-built structures).
  • Cache with unbounded insert and no eviction—semantic leak (still reachable) but OOM in practice.
  • Singleton holding shared_ptr to graphs that are never released at shutdown; still reachable in Valgrind, memory not returned to OS depending on allocator.
  • Thread exits with detached background work still allocating into structures owned from main thread—lifetime bugs sometimes appear as “leak under stress.”
  • Container of raw pointers from third-party C APIs (malloc)—free in destructors, or wrap in a small RAII guard class.
  • Double ownership: two delete is UB; not a leak, but a paired bug class.

False positives: reachable vs lost

“False positive” is overloaded:

  1. Tool classification: Valgrind can report still reachable blocks at exit. That is not a false report—it is a true unfreed block—but you may not care if it is a global cache freed only on real shutdown (or never in long-lived services where you still want bounded memory).
  2. Still reachable can hide growing caches if the cache is unbounded; it is not “free memory.”
  3. Definitely lost is what you should almost always fix unless you are in a quarantine phase with a suppression for a third-party bug on a path you cannot change yet.
  4. LSan/Valgrind and globals: static objects destruct in an order; some leaks appear only when races and atexit ordering interact—reduce complexity in shutdown paths.

Indirect leak (Valgrind/LSan terminology): blocks that are only referenced from other leaked blocks. When the root object of a structure is lost, its children still point to each other, so the tool reports the root as a direct (definite) leak and the children as indirect. The fix is almost always at the root object’s owner.


Suppression files

Suppressions list allow-listed frames and leak types for Valgrind, so CI can stay green for known third-party issues while your code is fixed forward.

Valgrind suppression file example (syntax illustrative; see Valgrind documentation for the exact memcheck:Leak or ... form your version needs):

{
   third-party-lib-ssl-init
   Memcheck:Leak
   fun:malloc
   fun:libSSL_internal_init_*
   ...
}

Workflow:

  1. Reproduce a leak with a known top stack in a foreign library.
  2. Guard upgrades of that library so suppressions are revisited on version bumps.
  3. Do not grow suppressions without an owner and a ticket to remove them.

LeakSanitizer also supports suppression lists via runtime options in several LLVM versions: check LSAN_OPTIONS=suppressions=/path/file in your clang release documentation.

Team policy tip: a suppression file in version control with comments and links to upstream bug trackers beats a silent exitcode=0 hack in CI that hides your own regressions.


CI integration

A typical CI matrix (Linux job shown):

# Example: build with ASan+LSan, run tests, fail on leak
c++ -std=c++20 -g -O1 -fsanitize=address -fno-omit-frame-pointer \
    -c test_main.cpp
c++ -fsanitize=address -fno-omit-frame-pointer test_main.o -o test_runner
ASAN_OPTIONS=detect_leaks=1:halt_on_error=1 \
    ./test_runner --gtest_color=yes

Valgrind in CI is often a nightly or merge-queue job because of speed:

valgrind --error-exitcode=1 --leak-check=full --errors-for-leak-kinds=definite \
    ./test_runner

Trade-offs: ASan/LSan jobs need sufficient RAM; Valgrind may need a bigger timeout multiplier (×10–50 depending on the binary).

Coverage: if a branch never runs in tests, a leak there will not be reported. Pair leak checks with good path coverage and fuzzing for parsers.

CMake: sanitizer presets

A CMake INTERFACE library keeps flags consistent for every test target that opts in:

# Example sketch — adjust generator expressions for multi-config
add_library(sanitize_address INTERFACE)
target_compile_options(sanitize_address INTERFACE
  $<$<CXX_COMPILER_ID:GNU,Clang>:-fsanitize=address -fno-omit-frame-pointer -g -O1>)
target_link_options(sanitize_address INTERFACE
  $<$<CXX_COMPILER_ID:GNU,Clang>:-fsanitize=address>)
# ...then: target_link_libraries(my_tests PRIVATE sanitize_address)

MSVC uses different flag spellings; guard with MSVC and the /fsanitize=address family (verify VS version). Keep separate targets for ASan and UBSan/TSan when mutually exclusive on your compiler version.

Ninja + ccache makes sanitized rebuilds tolerable; Bazel’s --config=asan pattern is the same idea: one place for cflags and link flags.

GitHub Actions: minimal Linux job

# .github/workflows/cpp-asan-lsan.yml (illustrative)
jobs:
  asan:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Configure and test with ASan+LSan
        env:
          ASAN_OPTIONS: detect_leaks=1:abort_on_error=1:print_stats=1
        run: |
          cmake -B build -DCMAKE_BUILD_TYPE=Debug -DENABLE_SANITIZERS=ON
          cmake --build build -j
          ctest --test-dir build --output-on-failure

Set CTEST_OUTPUT_ON_FAILURE to 1 in GitLab or other CI for parity. Double the default timeout for Valgrind jobs compared to ASan jobs on the same test suite.

Docker (Linux CI): install Valgrind and (optionally) heaptrack in the image. Tune ulimit (open file descriptors, address space) for integration tests that mmap large files; otherwise failures can look like memory pressure but are really resource limits.

Bazel (sketch): a --config=asan line in .bazelrc is the Bazel idiom parallel to CMake interface libraries—example:

# .bazelrc
build:asan --copt=-fsanitize=address --copt=-fno-omit-frame-pointer --copt=-g
build:asan --linkopt=-fsanitize=address
build:asan --test_env=ASAN_OPTIONS=detect_leaks=1

Run with bazel test --config=asan //... on a subset first; full tree can exhaust RAM on shared runners—shard or use tags like leak=heavy.

Static analysis (clang-tidy, compiler warnings)

-Wall -Wextra -Wconversion and clang-tidy checks such as cppcoreguidelines-*, clang-analyzer, and bugprone-* do not find all heap leaks, but they do find dangling writes, mismatched ownership, and bad moves that become leaks after refactors. Run tidy on the same translation units you ASan- test.


Production monitoring

The tools above are for development and CI. For a live system, the usual approach is to:

  • Track RSS and private bytes (Windows) per process and per service tier.
  • Correlate memory growth with restarts, load, and deploy events.
  • Use low-overhead heap profiling where sanitizers cannot run, such as jemalloc or tcmalloc heap profiles, or eBPF tools like bcc’s memleak, which record allocation stacks of outstanding blocks in a running process.

Caution: jemalloc / mimalloc and glibc may not return memory to the OS as aggressively; RSS can be flat or grow without a C++-level “leak” in the LSan sense—distinguish fragmentation and allocator caching from unreachable blocks using dev-time profilers. A/B a suspected build in staging with the same allocator as prod.

OOM: when the process dies, have core or at least logs and metrics; post-mortem without ASan/Valgrind means reproducing in test is mandatory.


Real examples

Example A: exception-unsafe resource release

Problem: a file descriptor and a heap block are interleaved; an exception in read_payload() skips cleanup.

Symptom: LSan: direct leak, stack at new in read_payload path.

Fix extract:

void handle_stream(std::istream& in) {
    std::string header;
    in >> header; // may throw on bad state
    auto body = std::make_unique<std::vector<char>>(read_size_from(header));
    read_payload(in, *body);
    process(*body);
}

Pair I/O and heap with RAII and one owning abstraction per level.

Example B: graph with shared_ptr cycle

Problem: a document model uses shared_ptr for parent and child both directions; closing the document “should” free everything, but reference count never hits zero.

Fix: one direction is weak_ptr<Parent> from Child, or a central Document owner and raw non-owning child pointers, depending on invariants. See the dedicated circular reference post.

Example C: third-party C API

Problem: you call C functions returning malloc’d buffers. Forgetting the documented free is a classic leak. Wrapping in a small struct with custom deleter on unique_ptr or a unique_ptr with a lambda deleter contains the rule in one place.

#include <cstdlib>
#include <memory>

extern "C" void* c_api_get_buffer(size_t* out_len);

struct FreeDeleter {
    void operator()(void* p) const noexcept { std::free(p); }
};

void use_c() {
    size_t n = 0;
    std::unique_ptr<void, FreeDeleter> p(c_api_get_buffer(&n));
    (void)n;
    // use p.get()
}

A function-object deleter costs nothing in the unique_ptr’s size, whereas a function-pointer deleter adds a pointer. Taking the address of standard library functions such as std::free is also not guaranteed to be portable in C++20, which is another reason to wrap the call.

Example D: callback capturing shared_ptr to a parent

Problem: a timer or network callback uses std::function (or a C-style void* user cookie) and captures shared_ptr<Session> by value to keep the session alive. If the callback is re-registered on every request and old registrations are not removed, you retain more Session state than you intended. In metrics, “open session” counts may climb on reconnect storms; a leak at process exit can be hard to see if a global list still holds one strong reference.

Fix direction: weak_ptr in the callback, ID-based session lookup in a central registry with erase-on-close, and explicit API contracts for unregistration. Add a stress test that connect/disconnects in a loop under ASan+LSan.

Example E: per-thread arena without teardown

Problem: a thread_local arena (for example a std::vector of scratch buffers) grows during request handling but is only cleared on the success path. Aborted or error paths leave the arena unbounded. The memory is “still reachable” from thread-local storage, so some tools will not label it a classic leak, but the process’s RSS can grow without limit.

Fixes: a small RAII guard that clear()s or pops the arena on scope exit (scope_exit from the Library Fundamentals TS v3 as std::experimental::scope_exit, or an equivalent helper, since it is not part of C++20), and a code-review rule that every control-flow exit from a handler resets thread-local scratch state.


Testing for leaks

Unit test patterns:

  • Build a shared asan config in CMake or Bazel, and link it only for asan test targets.
  • For hermetic cases, a test fixture that runs a function N times in a loop and asserts RSS (fragile) is worse than LSan. Prefer LSan/Valgrind in CI.
  • GoogleTest and others can fork; ensure leak checks are configured for the child in Valgrind if tests spawn.
  • For allocators with tracking_allocator, assert stats.live == 0 at scope exit (white-box).

Fuzzing (libFuzzer + ASan) finds leaks in parsers with inputs that are hard to hand-write. Run a corpus in CI to prevent regressions.

GoogleTest: death tests, forks, and LSan

EXPECT_DEATH, forking harnesses, and out-of-process isolations complicate exit-time leak scanning because the leaking process might be a child with its own heap. Run a subset of tests in a single process under ASan+LSan; use GTEST filters to separate “must be in-process for LSan” from “spawn only when needed.” For Valgrind, --trace-children=yes is often necessary; document it in your README for contributors.

Catch2, doctest, and custom main

Many frameworks define main() or wrap it: ensure ASan+LSan still run the fini path—return from main, not _exit, in tests (unless you intentionally test abnormal exit). std::quick_exit skips destructors; leak checkers at process end may mis-classify or under-report.

Regression: golden allocation counts (use sparingly)

White-box tests can assert that a function does not allocate in a steady state (e.g. reusing a buffer) by hooking operator new in test builds only. This is brittle across stdlib changes; prefer ASan+LSan and module-level invariants first.

Integration tests: run a scripted scenario and assert RSS delta is under a threshold in staging only—noisy, but can correlate with user-visible leaks when tools in dev are green.


Performance impact

The AddressSanitizer documentation describes a typical slowdown of about 2x, and memory use grows by a few times because of shadow memory and quarantine for freed blocks. Valgrind’s memcheck runs the program on a synthetic CPU and is typically one to two orders of magnitude slower, but it needs no special build, so it can check a release-shaped binary built with symbols. Heaptrack and Massif sit in between; they are usually much faster than memcheck and answer a different question: where growth comes from over time, rather than what is still allocated at exit.

A common split is to run ASan+LSan on every pull request or nightly, use Valgrind as a second opinion or for binaries that cannot be rebuilt, and use heaptrack or Massif when memory grows steadily in production and you need time on the x-axis.


Ownership habits that prevent leaks

Default ownership to std::unique_ptr, std::vector, and std::string, and use std::shared_ptr only when the design really shares lifetime.

For every C API that returns a malloced buffer, encode the free rule once, in a unique_ptr with a deleter or a small guard type, and document who frees it.

Treat exceptions and early returns as hostile to raw new: if a scope has more than one exit, move the allocation into a unique_ptr or a container.

In graphs and UI-style back-pointers, break shared_ptr cycles with weak_ptr or centralize ownership in one object.

In review, look closely at std::vector<T*> without a documented owner, at public APIs that return an owning raw pointer, and at static caches with no bound.

When RSS climbs in production but LSan is quiet, check for fragmentation, allocator caching, unbounded but reachable caches, and memory that never went through new (GPU memory, mmap) before concluding there is no leak.


Troubleshooting workflow

  1. Reproduce the growth or exercise the suspected path in a development build with -g and ASan+LSan, with a single thread if possible for cleaner stacks.
  2. When LSan reports a leak, read the allocation stack, fix the owner, rebuild, and re-run. For indirect leaks, walk up to the root object that should have been destroyed.
  3. When LSan is clean but Valgrind shows still reachable blocks, decide whether they are one-time initialization or a structure that grows per user or per request.
  4. When tools disagree with observed RSS, run heaptrack or Massif on a workload that matches production and map the growth to stacks, queues, caches, and per-connection state.
  5. For problems that appear only under load, also run a ThreadSanitizer build; data races can look like strange lifetime behavior.
  6. Add a regression test that would have failed LSan before the fix. If the leak depends on specific input, add that input to the test set, and for parsers, add fuzzing.

Platform differences: Linux, Windows, macOS

  • Linux: Valgrind, ASan/LSan, heaptrack, and perf are all available and are the most complete toolset.
  • Windows: Visual Studio heap snapshots and the CRT debug heap report leaks; MSVC’s AddressSanitizer catches corruption but not leaks; WSL2 can run Linux tools on Linux builds.
  • macOS: ASan from Clang for memory errors, and Xcode Instruments (Leaks, Allocations) for leak detection with a GUI. Valgrind support on recent macOS versions is limited.
  • Cross-platform projects: run one sanitizer job per OS, and use the same allocator in staging as in production when interpreting RSS.

Windows without WSL: MSVC, ASan, and Dr. Memory

Without a Linux build, Windows-native leak hunting relies on Visual Studio heap snapshots (take two snapshots around a repeated operation and diff them) and the Debug CRT leak report at exit. An ASan-enabled CI configuration still helps catch the memory corruption bugs that often accompany ownership mistakes.

Dr. Memory runs on Windows and reports leaks with call stacks without recompiling. As with any single tool, confirm noisy stacks with a second signal, especially in optimized builds where inlining blurs frames.

WSL2 and “Linux tools for Windows” teams

WSL2 runs a real Linux kernel. If production is also Linux, running Valgrind or heaptrack on a glibc build inside WSL2 is a close match to server behavior. Keep glibc and compiler versions roughly aligned with production, because malloc behavior can differ between distributions.

macOS: ASan vs Instruments

Command-line ASan is easy to script in CI. Xcode Instruments (Leaks, Allocations) adds context for AppKit run loops and GCD queues. For portable C++ libraries, reproduce the issue in a headless test first, and use Instruments when the leak only shows up inside a GUI or system framework.


Classifying tool output (Valgrind/LSan cheat sheet)

Definitely lost (Valgrind) or direct leak (LSan): no pointer to the block remains. Start at the allocation stack and give that block a real owner.

Indirectly lost or indirect leak: the block is reachable only from other leaked blocks. Usually a container or parent object leaked; fixing child pointers does not help until the root owner is fixed.

Possibly lost (Valgrind): only an interior pointer into the block remains, not a pointer to its start. This can be a real leak, or a structure that intentionally keeps offset pointers (some string implementations and custom allocators do).

Still reachable: a pointer still exists at exit. It can be harmless one-time initialization or a cache that grows with every user; if the reachable set grows without a bound, treat it as a production leak regardless of the label.

When LSan prints <unknown module> frames, rebuild with -g, check for stale symbols, and make sure llvm-symbolizer is on the PATH, especially for code loaded with dlopen and unloaded before exit.

On MSVC Debug builds, the CRT’s end-of-process heap dump lists block numbers; defining _CRTDBG_MAP_ALLOC adds file and line for malloc-family calls, and _CrtSetBreakAlloc(n) breaks in the debugger when block number n is allocated on the next run.


Conclusion

Memory leaks in C++ are ownership and lifetime bugs: a heap block can no longer be freed under the program’s own rules, so it survives until the process ends, sometimes slowly enough to hurt only long-running systems. The prevention is to give every allocation a clear owner (unique_ptr, value-semantic containers, and shared_ptr graphs with weak_ptr on back-edges), and the safety net is ASan+LSan in development and CI, Valgrind for depth or unmodified binaries, and heap profilers when steady growth is the symptom. Suppression files should document known third-party issues, not hide new regressions, and production monitoring covers the paths that CI never stresses.

Further reading on this blog: detection in practice, Valgrind, circular shared_ptr issues, and the series post on real leak patterns for narrative-style case studies.