C++ Memory Leak Debugging Case Study

Key takeaways

A chat server's memory grows steadily until the OOM killer steps in. The walkthrough narrows it down with metrics, ASan's LeakSanitizer and Heaptrack to event listeners that are never unsubscribed, fixes it with RAII subscriptions, and adds sanitizers to CI.

Introduction

Memory leaks in long-running servers rarely crash anything right away. They show up as a slow, steady rise in memory until the process is restarted or killed. This article walks through one such leak, from first symptoms through root cause, fix, and prevention.

The scenario is a composite built around one of the most common leak patterns in event-driven C++ code: callbacks that are registered and never removed. The command outputs below are illustrative and trimmed to the relevant lines, so treat the numbers as examples of the shape of the data, not as measurements.

What this walkthrough covers

  • How to recognize leak-like memory growth early
  • When to use Valgrind, ASan, and Heaptrack, and what each can and cannot see
  • How to narrow a leak down in a codebase you did not write
  • Coding patterns that prevent this class of leak

Symptom: server memory keeps growing

What we saw

The service is a chat server. A few days after a deploy, memory use is clearly higher than at start and keeps rising roughly linearly.

# Right after deploy
$ ps aux | grep chat_server
user  12345  0.5  2.1  524288  ...  ./chat_server
# Three days later
$ ps aux | grep chat_server
user  12345  0.5  8.7  2162688  ...  ./chat_server
# Seven days later (killed by OOM killer)
[  123.456] Out of memory: Killed process 12345 (chat_server)

The OOM line comes from the kernel log (dmesg or journalctl -k), not from the application, which is why a service that is “restarted by the supervisor every few days” can hide a leak for a long time. The fourth column of ps aux is %MEM and the sixth is RSS in kilobytes; RSS is the number to watch, because virtual size (VSZ) also counts reserved but untouched address space and thread stacks.

Early hypotheses

  1. Are connection objects not freed properly?
  2. Is a log buffer growing without bound?
  3. Is a cache growing forever?

Writing hypotheses down before opening a profiler matters, because each one predicts a different correlation. Per-connection leaks track the number of connections; buffers track the volume of traffic; caches track the number of distinct keys. The metrics in the next step are chosen to tell these apart.


Initial analysis: monitoring data

Prometheus metrics

// Metrics collection added to the server
class MemoryMetrics {
public:
    static size_t getCurrentRSS() {
        std::ifstream stat("/proc/self/status");
        std::string line;
        while (std::getline(stat, line)) {
            if (line.find("VmRSS:") == 0) {
                std::istringstream iss(line);
                std::string key, value, unit;
                iss >> key >> value >> unit;
                return std::stoull(value) * 1024; // KB to bytes
            }
        }
        return 0;
    }
};
// Send metrics periodically
void reportMetrics() {
    auto rss = MemoryMetrics::getCurrentRSS();
    prometheus_gauge_set(memory_rss_bytes, rss);
}

Reading /proc/self/status is Linux-specific and cheap enough to do every few seconds. If you use the official Prometheus client libraries, a process collector usually exports process_resident_memory_bytes for you, which is the same value.

Pattern

From the dashboard:

  • Memory: grows at a steady rate, day and night
  • Connection count: stable (100–200)
  • Throughput: unchanged

Conclusion: it is not “memory per open connection” but something that accumulates over time. The growth rate is a better clue than the absolute value. Growth that continues at night when traffic is low points away from buffers and toward something tied to events that keep happening, such as users joining and leaving.


Tool choice: Valgrind vs ASan vs Heaptrack

Comparison

ToolStrengthsWeaknessesBest for
Valgrind (Memcheck)No recompile needed, precise leak reportsVery slow (the manual cites 10–50×)Dev, small repro cases
ASan + LeakSanitizerSmall slowdown (typically around 2×), many bug classesNeeds recompile; leak check only at exitCI, integration and load tests
HeaptrackAllocation timeline and call-path attributionOverhead grows with allocation rateMemory growth profiling

Strategy

  1. Try ASan for quick reproduction under realistic load
  2. If it does not reproduce, use Valgrind on a narrowed-down case
  3. Use Heaptrack to see where memory is held while the program runs

The key limitation to keep in mind is what “leak” means to each tool. Valgrind and LSan report memory that is unreachable when the program ends. Memory that is still referenced from a live container at exit is freed or reported as “still reachable”, even if that container grew without bound. Heaptrack, by contrast, records every allocation and free with its call stack, so it can show memory that is technically reachable but should have been released long ago.


First pass with Valgrind

Build and run

# Debug symbols, no optimization
$ g++ -g -O0 -std=c++17 *.cpp -o chat_server
# Run under Valgrind
$ valgrind --leak-check=full --show-leak-kinds=all \
           --track-origins=yes --log-file=valgrind.log \
           ./chat_server

Problem

The server became too slow to handle anything close to real load, so in a ten-minute run the growth was too small to stand out.

==12345== HEAP SUMMARY:
==12345==     in use at exit: 1,234,567 bytes in 1,234 blocks
==12345==   total heap usage: 12,345 allocs, 11,111 frees, 123,456,789 bytes allocated

Conclusion: Valgrind is too slow to replay production-like load here. --track-origins=yes makes it slower still; it helps with uninitialized-value errors and adds nothing for leak hunting, so it can be dropped in a run like this.


Fast reproduction with ASan

ASan build

# Recompile with ASan
$ g++ -g -O1 -fsanitize=address -fno-omit-frame-pointer \
      -std=c++17 *.cpp -o chat_server_asan
$ export ASAN_OPTIONS=detect_leaks=1:log_path=asan.log

-fno-omit-frame-pointer keeps stack traces reliable with ASan’s fast unwinder, and -O1 keeps the slowdown moderate while leaving the code debuggable. On Linux x86-64 detect_leaks=1 is already the default; on macOS LeakSanitizer support depends on the toolchain.

Load test

# Simulate real traffic
$ ./load_test.sh --connections=200 --duration=600s

The load test must exercise the same lifecycle as production, in this case users repeatedly joining and leaving rooms, not only sending messages over long-lived connections. A test that opens 200 connections and keeps them open would never trigger this leak.

Result

After the load test, shutting the server down cleanly produced this report:

=================================================================
==23456==ERROR: LeakSanitizer: detected memory leaks
Direct leak of 48000 byte(s) in 1000 object(s) allocated from:
    #0 0x7f123456 in operator new(unsigned long)
    #1 0x7f234567 in EventManager::subscribe(std::string const&, EventCallback)
    #2 0x7f345678 in ChatRoom::addUser(User*)
    #3 0x7f456789 in Server::handleJoin(Connection*)
    ...
SUMMARY: AddressSanitizer: 48000 byte(s) leaked in 1000 allocations.

Finding: the leaked blocks are allocated in EventManager::subscribe, called from ChatRoom::addUser. Note the word clean: LSan runs its check at normal process exit. A process killed with SIGKILL, or one that calls _exit, never reports anything, so the server needs a shutdown path that returns from main for this technique to work.


Allocation patterns with Heaptrack

Running Heaptrack

$ heaptrack ./chat_server
$ heaptrack_gui heaptrack.chat_server.12345.gz

Heaptrack can also attach to a running process with heaptrack -p <pid>, which is useful when the problem only appears after hours of uptime. heaptrack_print gives a text summary when no GUI is available.

Findings

From the flame graph and the “consumed memory over time” chart:

  1. EventManager::subscribe is one of the largest sources of memory still held at the end
  2. The allocations from that call path keep growing, with almost no matching frees
  3. Call stack: ChatRoom::addUser → subscribe

This confirms that the LSan report is the real growth and not a one-off leak at shutdown: memory from this path climbs throughout the run.


Root cause: accumulating event listeners

Buggy code

class EventManager {
    std::unordered_map<std::string, std::vector<EventCallback*>> listeners_;
public:
    void subscribe(const std::string& event, EventCallback callback) {
        // Bug: allocated with new but never freed
        auto* cb = new EventCallback(std::move(callback));
        listeners_[event].push_back(cb);
    }
    
    void publish(const std::string& event, const EventData& data) {
        if (auto it = listeners_.find(event); it != listeners_.end()) {
            for (auto* cb : it->second) {
                (*cb)(data);
            }
        }
    }
    
    // Destructor does not free listeners!
    ~EventManager() = default;
};
class ChatRoom {
    EventManager& eventMgr_;
    
public:
    void addUser(User* user) {
        // Register a listener on every join
        eventMgr_.subscribe("message", [user](const EventData& data) {
            user->sendMessage(data);
        });
        
        // User leaves but listeners remain!
    }
};

Why it leaked

  1. Every addUser did new EventCallback
  2. After a user left, pointers stayed in listeners_
  3. The destructor did not free them
  4. N joins → N allocations → 0 frees, and the vector of listeners grows with every join

There are really two bugs here, and it is important to separate them. The raw new without delete is what LeakSanitizer reports, but it only becomes visible at shutdown. The bug that makes memory grow while the server runs is the missing unsubscribe: even with perfect ownership, a vector that gains one callback per join and never loses one is unbounded growth. It is also a correctness bug. Each stale callback captured a raw User*; once that user is destroyed, the next publish("message", ...) calls sendMessage on freed memory. In practice this kind of code often leaks for weeks and then starts crashing intermittently in publish, and the crash is the symptom people notice first.


Fix: RAII and smart pointers

Option 1: smart pointers

class EventManager {
    using CallbackPtr = std::shared_ptr<EventCallback>;
    std::unordered_map<std::string, std::vector<CallbackPtr>> listeners_;
public:
    // Returns subscription id for later unsubscribe
    size_t subscribe(const std::string& event, EventCallback callback) {
        auto cb = std::make_shared<EventCallback>(std::move(callback));
        listeners_[event].push_back(cb);
        return reinterpret_cast<size_t>(cb.get());
    }
    
    void unsubscribe(const std::string& event, size_t id) {
        auto& cbs = listeners_[event];
        cbs.erase(
            std::remove_if(cbs.begin(), cbs.end(),
                [id](const CallbackPtr& cb) {
                    return reinterpret_cast<size_t>(cb.get()) == id;
                }),
            cbs.end()
        );
    }
    
    ~EventManager() = default; // shared_ptr cleans up
};

Smart pointers fix the ownership bug: nothing leaks at shutdown any more. On their own they do not fix the growth, because the vector still holds every callback until someone calls unsubscribe. Two details in this version are worth improving before shipping it. Using the callback’s address as the ID can collide once a callback is freed and the allocator reuses its address for a new one, so an old ID could remove the wrong subscription; a monotonically increasing counter avoids that. And listeners_[event] inside unsubscribe inserts an empty entry for unknown events; find is the safer lookup. A plain std::unique_ptr (or storing EventCallback by value) would be enough here, since nothing shares ownership of the callbacks.

Option 2: RAII wrapper

class Subscription {
    EventManager* mgr_;
    std::string event_;
    size_t id_;
public:
    Subscription(EventManager* mgr, std::string event, size_t id)
        : mgr_(mgr), event_(std::move(event)), id_(id) {}
    
    ~Subscription() {
        if (mgr_) {
            mgr_->unsubscribe(event_, id_);
        }
    }
    
    Subscription(Subscription&& other) noexcept
        : mgr_(other.mgr_), event_(std::move(other.event_)), id_(other.id_) {
        other.mgr_ = nullptr;
    }
    
    Subscription(const Subscription&) = delete;
    Subscription& operator=(const Subscription&) = delete;
};
class ChatRoom {
    EventManager& eventMgr_;
    std::vector<Subscription> subscriptions_;
public:
    void addUser(User* user) {
        auto id = eventMgr_.subscribe("message", [user](const EventData& data) {
            user->sendMessage(data);
        });
        
        subscriptions_.emplace_back(&eventMgr_, "message", id);
    }
    
    void removeUser(User* user) {
        // Removing from subscriptions_ triggers unsubscribe
        // (in practice, map users to subscriptions)
    }
};

This is the fix that addresses the root cause, because it ties the listener’s lifetime to an object whose lifetime you already manage. When the subscription is destroyed, the callback is removed, so a user who leaves can no longer be called. The natural next step is to store the Subscription inside the per-user state (for example std::unordered_map<User*, Subscription>), so destroying a user’s entry unsubscribes it automatically.

Two C++ details matter here. Subscription declares a move constructor but no move assignment, and its copy assignment is deleted, so it is not move-assignable. std::vector::erase needs move assignment to shift the remaining elements, which means erasing from subscriptions_ will not compile until you add Subscription& operator=(Subscription&&) noexcept (unsubscribe the current one, then take over the other’s state). The move constructor is noexcept, which lets vector move rather than copy on reallocation; since copying is deleted, it could not reallocate otherwise. Second, the wrapper holds a raw EventManager*, so the manager must outlive every subscription; declaring the manager before the rooms that use it, so it is destroyed after them, is the simplest way to guarantee that.


Verification: comparing memory profiles

Before

$ heaptrack ./chat_server_before
# After 10 minutes of the join/leave load test
peak heap memory consumption: grows with the number of joins
total memory leaked: dominated by EventManager::subscribe

After

$ heaptrack ./chat_server_after
# Same load test
peak heap memory consumption: flat after warm-up
total memory leaked: only small one-time allocations (static objects)

The comparison that matters is the shape of the “consumed memory over time” curve under the same load: before the fix it rises steadily, after the fix it levels off once caches and pools have warmed up. A flat heap curve with a slowly rising RSS usually points to allocator behavior or fragmentation rather than another leak.

ASan final check

$ ./chat_server_asan
# After 10 min load test, clean exit: LeakSanitizer prints nothing when no leaks are found
$ echo $?
0

When no leaks are found, LeakSanitizer is silent and the exit code is unchanged. When it does find leaks, it prints the report and exits with a non-zero status (configurable with the exitcode= option), which is what makes it usable as a CI gate.


Prevention: ASan in CI

GitHub Actions

# .github/workflows/sanitizers.yml
name: Memory Sanitizers
on: [push, pull_request]
jobs:
  asan:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      
      - name: Build with ASan
        run: |
          cmake -DCMAKE_BUILD_TYPE=Debug \
                -DCMAKE_CXX_FLAGS="-fsanitize=address -fno-omit-frame-pointer" \
                -B build
          cmake --build build
      
      - name: Run tests with ASan
        run: |
          export ASAN_OPTIONS=detect_leaks=1:halt_on_error=1
          cd build && ctest --output-on-failure

CI with sanitizers catches ownership leaks in any code path the tests execute, so it is only as good as the tests. For this particular bug, a unit test that joins and removes a user and then asserts that the event manager has zero listeners would have caught the missing unsubscribe as well, which no sanitizer can.

Code review questions

  • If new is used, who owns the result, and is it a smart pointer?
  • Is RAII used for resource acquisition?
  • For every callback or listener registration, where is the matching unregistration?
  • Does a callback capture a raw pointer or reference to an object that can be destroyed first?

Lessons and patterns

Takeaways

  1. Detect early: export memory metrics from the first deploy
  2. Combine tools: ASan for reproduction, Heaptrack for growth over time, Valgrind when you cannot rebuild
  3. RAII: tie every registration to an object whose destructor undoes it
  4. Automate: sanitizers in CI to catch regressions

Patterns that help avoid leaks

// Bad: manual memory
class BadCache {
    std::map<std::string, Data*> cache_;
public:
    void add(const std::string& key, Data* data) {
        cache_[key] = data; // who deletes?
    }
};
// Good: smart pointers
class GoodCache {
    std::map<std::string, std::unique_ptr<Data>> cache_;
public:
    void add(const std::string& key, std::unique_ptr<Data> data) {
        cache_[key] = std::move(data);
    }
};
// Better: value semantics
class BestCache {
    std::map<std::string, Data> cache_;
public:
    void add(const std::string& key, Data data) {
        cache_[key] = std::move(data);
    }
};

All three caches, including the “best” one, share the property that caused this incident: nothing ever removes entries. Value semantics solve ownership; they do not bound growth. A cache in a long-running server needs an eviction policy (size limit, LRU, TTL) just as a listener list needs an unsubscribe path.


Closing thoughts

  • Leaks often show up slowly, so memory monitoring is part of running a server, not an optional extra
  • Know what each tool sees: leak checkers find unreachable memory at exit, heap profilers find growth while running
  • RAII and smart pointers fix ownership; lifecycle design (unsubscribe, eviction) fixes growth
  • CI sanitizers and lifecycle tests catch regressions early

FAQ

Q1. Can we run ASan in production? Its slowdown and extra memory use (shadow memory and quarantine) make it unusual for full production traffic. Common alternatives are routing a small fraction of traffic to an ASan build, or replaying captured production traffic against an ASan build in staging.

Q2. Valgrind says “still reachable”—is that a leak? “Still reachable” means memory still pointed to at exit. That is normal for static singletons; if the amount grows with runtime or load, it is exactly the kind of unbounded growth described in this article.

Q3. Don’t smart pointers cause leaks via cycles? They can: two objects holding shared_ptr to each other are never freed. Break the cycle with weak_ptr, and prefer unique_ptr when ownership is clear. See circular references with shared_ptr.