C++ Intermittent Segfault Case Study: Core Dumps, gdb, and rr

Key takeaways

An intermittent segfault in a multithreaded C++ server, traced from a core dump to a check-then-use data race with gdb and rr, then fixed and verified with ThreadSanitizer.

Introduction

“The server sometimes dies” is one of the least useful bug reports and one of the most common in multithreaded C++. This case study walks through an intermittent segfault in a chat server built on Boost.Asio with several worker threads. The crash turned out to be a classic check-then-use data race on a raw pointer. The bug itself is ordinary. What is worth studying is the process: each tool answers one specific question, and knowing which question to ask next is most of the work.

  • The core dump answers: where did it fail, and what did memory look like at that instant?
  • gdb on the core answers: which thread, which call chain, which object?
  • rr answers the question a core dump cannot: who changed that state, and when?
  • ThreadSanitizer answers: is this really a race, and is the fix complete?

Symptom: intermittent SIGSEGV

The kernel log showed a crash every day or two, always under load:

$ dmesg | tail
chat_server[23456]: segfault at 0 ip 00007f1234567890 sp 00007fff12345678 error 4 in chat_server

segfault at 0 means the faulting address was 0 (or very close to it), which points to a null pointer dereference. error 4 is the page-fault error code: a user-mode read of a page that is not present. A read at a small address such as 0x8 or 0x18 usually means “null object pointer plus a member offset”, which is a useful hint even before you open a debugger.

The local load test did not reproduce it. That already says something. A deterministic logic bug usually shows up under a load test. A crash that depends on production traffic patterns and CPU count is most likely timing-dependent.

Getting a core dump

$ ulimit -c unlimited
$ sudo sysctl -w kernel.core_pattern=/var/coredumps/core.%e.%p.%t
$ sudo mkdir -p /var/coredumps
$ sudo chmod 1777 /var/coredumps

This setup has a few traps. ulimit -c applies only to the shell and its children. A service started by systemd ignores it and uses LimitCORE=infinity in the unit file instead. On many distributions core_pattern is already a pipe to systemd-coredump, so the dumps end up in the journal, not in a directory. In that case coredumpctl list and coredumpctl gdb <pid> are the right commands. core_pattern is also global to the host kernel, so inside a container you cannot set it, and whatever the host configured applies.

The core is only useful with the exact binary that crashed and its debug information. I have lost time loading a core against a binary rebuilt “from the same commit” with different flags, which gives a backtrace that looks plausible and is wrong. Keep the deployed binary, or at least its separate debug symbols (objcopy --only-keep-debug), for every release.

The next morning:

$ ls -lh /var/coredumps/
-rw------- 1 user user 1.2G Mar 30 03:42 core.chat_server.23456.1711756920

Reading the core in gdb

$ gdb ./chat_server /var/coredumps/core.chat_server.23456.1711756920
(gdb) bt
#0  ChatRoom::broadcast (this=0x0, msg=...) at src/chat_room.cpp:145
#1  Connection::handleMessage (this=0x55d0c3a8e2f0, msg=...) at src/connection.cpp:45
#2  ... asio handler frames ...
(gdb) info threads

this=0x0 in frame 0 means broadcast was called on a null ChatRoom*. Frame 1 is the interesting part, because handleMessage checks for null before calling:

class Connection {
    ChatRoom* room_;  // raw pointer, written by other threads

public:
    void handleMessage(const std::string& msg) {
        if (room_) {
            room_->broadcast(msg);
        }
    }

    void leaveRoom() {
        room_ = nullptr;
    }
};

A null this right after a null check has two explanations. Either the pointer was changed between the check and the call, or the check and the call read different values. Both are only possible if another thread writes room_ concurrently. info threads and thread apply all bt show what the other threads were doing at the moment of the crash, but by then the writer has usually finished and moved on. The core shows the victim, not the culprit.

Why the local repro failed

$ ./load_test.sh --users=1000 --duration=600
# No crash

The window between if (room_) and room_->broadcast is a few instructions long. To hit it, a disconnect on one thread must land in exactly that gap while another thread handles a message for the same connection. Production had more cores, real clients that disconnect in the middle of sending, and much longer uptime. A synthetic test that connects, sends, and disconnects in a neat order almost never lines those events up.

Adding logging made it worse. Each std::cout or logger call takes a lock and changes timing, and a race that needs a narrow window can disappear under that kind of instrumentation.

Recording with rr

rr records a process’s execution, including every source of nondeterminism such as system call results, signals, and thread scheduling. It can then replay that same execution as many times as needed under gdb, including backwards.

$ sudo apt install rr
$ echo 1 | sudo tee /proc/sys/kernel/perf_event_paranoid
$ rr record --chaos ./chat_server
rr: Saving execution to trace directory `/root/.local/share/rr/chat_server-0'.

rr has real constraints. It needs hardware performance counters, so it fails in many cloud VMs and containers that do not expose the PMU. It also runs all threads on a single core, which by itself makes many races far less likely. The --chaos flag counters that by randomizing scheduling decisions. In practice, I run the load test against rr record --chaos in a loop and delete the traces of runs that did not crash. When one run finally crashes, you have a trace you can replay as many times as you want.

Reverse debugging to the culprit

$ rr replay /root/.local/share/rr/chat_server-0
(rr) continue
Program received signal SIGSEGV, Segmentation fault.
ChatRoom::broadcast (this=0x0, msg=...) at src/chat_room.cpp:145
(rr) up
#1  Connection::handleMessage (this=0x55d0c3a8e2f0, msg=...) at src/connection.cpp:45
(rr) watch -l room_
(rr) reverse-continue
Hardware watchpoint 1: -location room_
Old value = (ChatRoom *) 0x0
New value = (ChatRoom *) 0x55d0c3a91a40
(rr) bt
#0  Connection::leaveRoom () at src/connection.cpp:67
#1  ChatRoom::removeConnection (...) at src/chat_room.cpp:89
#2  Server::handleDisconnect (...) at src/server.cpp:123

watch -l (short for -location) watches the memory address rather than the expression, which matters because the expression room_ is only meaningful in the frame where this is valid. Running reverse-continue goes backwards until that memory was last written. In reverse, gdb reports the values in the direction you are moving. “Old” is the value at the later point (null), and “New” is the value from before the write (the valid room pointer). The backtrace at that point is the answer the core dump could not give: a disconnect handler on a different thread cleared room_ while handleMessage was between its check and its call.

Root cause: check-then-use race

Time  | Thread A (message)           | Thread B (disconnect)
------|------------------------------|-----------------------
t0    | if (room_)   // non-null     |
t1    |                              | room_ = nullptr;
t2    | room_->broadcast(msg);       |
      | reads room_ again → nullptr  |
      | SIGSEGV                      |

Formally, any unsynchronized concurrent write and read of the same non-atomic object is a data race, and the whole program has undefined behavior. The compiler is allowed to assume no other thread changes room_. It may load it once and reuse the value, or load it twice. That is why the same bug can crash in one build and not in another. A variant is often worse than a clean null crash: if the room is deleted instead of the pointer being cleared, thread A calls into freed memory and corrupts the heap, and the crash happens later somewhere unrelated.

Fixes and their trade-offs

Option A: mutex

class Connection {
    std::mutex roomMutex_;
    ChatRoom* room_ = nullptr;

public:
    void handleMessage(const std::string& msg) {
        std::lock_guard<std::mutex> lock(roomMutex_);
        if (room_) {
            room_->broadcast(msg);
        }
    }

    void leaveRoom() {
        std::lock_guard<std::mutex> lock(roomMutex_);
        room_ = nullptr;
    }
};

This is correct for the pointer. It still holds the lock while broadcast sends to every other connection, and if broadcast ever takes another connection’s roomMutex_ (or a room lock that leaveRoom callers already hold), you have a lock-order deadlock. A tighter version copies the pointer under the lock and calls outside it, but that only works if the room’s lifetime is guaranteed some other way.

Option B: shared ownership with weak_ptr

class Connection {
    std::mutex roomMutex_;
    std::weak_ptr<ChatRoom> room_;

public:
    void handleMessage(const std::string& msg) {
        std::shared_ptr<ChatRoom> room;
        {
            std::lock_guard<std::mutex> lock(roomMutex_);
            room = room_.lock();
        }
        if (room) {
            room->broadcast(msg);  // room stays alive for this call
        }
    }

    void leaveRoom() {
        std::lock_guard<std::mutex> lock(roomMutex_);
        room_.reset();
    }
};

This fixes the lifetime problem as well: once lock() returns a shared_ptr, the room cannot be destroyed during broadcast. A common mistake is to assume weak_ptr alone makes this thread-safe. The control block is thread-safe, but the weak_ptr object itself is not. room_.lock() on one thread and room_.reset() on another is still a data race. You need the mutex, or std::atomic<std::weak_ptr<T>> in C++20.

Option C: Asio strand

class Connection : public std::enable_shared_from_this<Connection> {
    boost::asio::strand<boost::asio::io_context::executor_type> strand_;
    ChatRoom* room_ = nullptr;

public:
    void handleMessage(std::string msg) {
        boost::asio::post(strand_, [self = shared_from_this(), msg = std::move(msg)] {
            if (self->room_) self->room_->broadcast(msg);
        });
    }

    void leaveRoom() {
        boost::asio::post(strand_, [self = shared_from_this()] {
            self->room_ = nullptr;
        });
    }
};

A strand guarantees that handlers posted to it never run concurrently, so room_ needs no lock at all, as long as every access goes through the strand. That condition is the weak point: one direct call from a non-strand thread brings the race back, and nothing in the type system stops it. Capturing shared_from_this() rather than raw this also matters. Otherwise the connection can be destroyed while a posted handler is still queued, which replaces one crash with another.

We chose the strand because the codebase was already built around per-connection strands, and it avoided adding a second locking scheme. In a codebase without Asio, Option B is the most robust.

Verifying with ThreadSanitizer

$ g++ -g -O1 -fsanitize=thread -std=c++17 *.cpp -o chat_server_tsan -pthread
$ ./chat_server_tsan
WARNING: ThreadSanitizer: data race (pid=12345)
  Write of size 8 at 0x7b0c00000010 by thread T2:
    #0 Connection::leaveRoom() src/connection.cpp:67
  Previous read of size 8 at 0x7b0c00000010 by thread T1:
    #0 Connection::handleMessage() src/connection.cpp:45
SUMMARY: ThreadSanitizer: data race src/connection.cpp:67 in Connection::leaveRoom()

Running the load test against the unfixed TSan build reported the race right away, even though the crash itself needed days. TSan does not need the bad interleaving to happen. It detects two accesses to the same address that are not ordered by any synchronization. After the fix, the same test ran clean. TSan cannot be combined with ASan in one binary (the runtimes conflict), so run them as separate CI jobs. Also expect TSan to slow the program considerably and use much more memory. It belongs in tests and staging, not in production.

The honest limit: TSan only reports races on code paths that actually ran. A clean TSan run on a test that never disconnects mid-message proves nothing about this bug. The test has to exercise the concurrent path.

What I would do differently

Looking back, the biggest mistake was trusting the null check. if (p) p->f() looks defensive, and in single-threaded code it is. With a pointer shared across threads, it only moves the crash elsewhere. The review question should not be “is it checked?” but “who else can write this, and what orders their write against my read?”

The second lesson is ordering the tools. Starting with the core dump was right. Hours of “add a log line and wait a day” were wasted before anyone tried TSan on the load test, which would have pointed at the same two lines in minutes.

Useful gdb commands for this kind of hunt:

(gdb) thread apply all bt          # every thread's stack in the core
(gdb) info threads
(gdb) frame 1
(gdb) print *this
(gdb) watch -l this->room_         # on a live process or under rr
(gdb) break Connection::handleMessage if room_ == 0