C++ Multithreading Crashes: Data Races, mutex, and atomic

Key takeaways

Fix intermittent multithreaded crashes: data races vs race conditions, std::mutex, atomics, false sharing basics, condition variables, and ThreadSanitizer (-fsanitize=thread).

Data races are UB (undefined behavior). Iterator misuse across threads overlaps with iterator invalidation.

Introduction: “Multithreaded code crashes sometimes”

The most common serious bug in concurrent C++ is a data race: unsynchronized access to shared mutable state.

int counter = 0;
void worker() {
    for (int i = 0; i < 1'000'000; ++i) {
        ++counter;  // data race
    }
}

Run worker on two threads and the final value is usually less than 2,000,000, and different on every run. ++counter looks like one operation but compiles to three: load the value into a register, add one, store it back. When two threads interleave those steps, both load the same old value, both add one, and one increment is lost. With optimizations on, the compiler may even keep counter in a register for the whole loop and store it once at the end, because nothing in single-threaded semantics forbids that. The counter is the harmless version. The same race on a std::vector, std::string or std::map corrupts internal pointers, and the crash appears later, in unrelated code.

This article covers:

  • Data races vs informal “race conditions”
  • std::mutex patterns
  • std::atomic basics
  • ThreadSanitizer
  • Ten common concurrency bugs with short examples

What is a data race?

A data race (in the standard sense) requires conflicting accesses without synchronization: two threads touch the same memory location, at least one of them writes, and nothing orders the two accesses. It is undefined behavior, not merely “a wrong value”.

Typical outcomes: torn reads/writes, crashes, or “lucky” passes.

A race condition is the broader, informal term for any bug where the result depends on timing. You can have a race condition with no data race at all. Every individual access may be protected by a mutex or be atomic, and the program can still be wrong because the sequence is not protected:

std::atomic<int> balance{100};
void withdraw(int amount) {
    if (balance >= amount) {      // check (atomic read)
        balance -= amount;        // act (atomic write) — another thread may have withdrawn in between
    }
}

No data race exists here, so ThreadSanitizer stays silent, yet two concurrent withdrawals of 80 can both pass the check and drive the balance negative. The fix is to make check-and-act one operation: a mutex around both lines, or a compare_exchange loop. Knowing which kind of bug you have matters, because tools find data races mechanically, while race conditions only show up through reasoning about invariants and through testing.

Undefined behavior also explains why “add a printf and the bug disappears” is so common. The extra I/O changes timing and inhibits some optimizations, which is enough to hide a race without fixing it.


mutex

#include <mutex>
int counter = 0;
std::mutex mtx;
void worker() {
    for (int i = 0; i < 1'000'000; ++i) {
        std::lock_guard<std::mutex> lock(mtx);
        ++counter;
    }
}

std::lock_guard locks in its constructor and unlocks in its destructor, so the mutex is released on every path out of the scope, including exceptions. Calling mtx.lock() and mtx.unlock() by hand works until someone adds an early return in between, and then the next thread to lock blocks forever.

This version is correct but slow: it takes and releases the lock two million times, and under contention the threads spend most of their time waiting for each other. For work like this, let each thread count in a local variable and add its total under the lock once at the end. The general rule is to keep the shared, locked part small and infrequent, not to lock around every tiny step.

Deadlock avoidance

Always acquire multiple mutexes in a consistent global order, or use std::scoped_lock (C++17) / std::lock to lock several mutexes atomically.

void transfer(Account& from, Account& to, int amount) {
    std::scoped_lock lock(from.mtx, to.mtx);  // deadlock-free for any argument order
    from.balance -= amount;
    to.balance += amount;
}

The classic deadlock is transfer(a, b) on one thread and transfer(b, a) on another, each locking its first argument and waiting for the second. scoped_lock with several mutexes uses a deadlock-avoidance algorithm, so the argument order no longer matters. A deadlock is at least easy to diagnose once it happens: attach a debugger to the hung process and look at where each thread’s stack is waiting (thread apply all bt in GDB).


Atomics

#include <atomic>
std::atomic<int> counter{0};
void worker() {
    for (int i = 0; i < 1'000'000; ++i) {
        ++counter;  // atomic RMW
    }
}

++ on a std::atomic<int> is a single read-modify-write instruction (lock xadd on x86), so no increment is lost. By default atomic operations use memory_order_seq_cst, the strongest and simplest ordering. For a pure statistics counter that nothing else depends on, counter.fetch_add(1, std::memory_order_relaxed) is enough and can be cheaper on ARM, but relaxed ordering must not be used for flags that publish other data: a “ready” flag needs release on the store and acquire on the load, or the reader may see the flag before the data it guards.

Atomics are not free under contention either. Every increment forces the cache line holding counter to move between cores, so many threads hammering one atomic can be slower than per-thread counters combined at the end.

Rule of thumb: one simple counter or flag → often atomic. Multiple fields that must move together → mutex. std::atomic<std::string> does not compile, and std::atomic of a large struct is usually implemented with a hidden lock anyway (check is_lock_free()).


ThreadSanitizer

g++ -g -fsanitize=thread -std=c++17 -o myapp main.cpp
./myapp

TSan reports racing lines with stacks—fix the synchronization at those sites. A report starts with WARNING: ThreadSanitizer: data race and shows two stacks: the current access (“Write of size 4 at … by thread T2”) and the previous conflicting one (“Previous write … by thread T1”), plus where each thread was created. Fix the synchronization at both sites; protecting only one side still leaves a race.

Some practical points. TSan instruments memory accesses at runtime, so programs typically run several times slower and use much more memory; it is a test-suite tool, not a production one. It only finds races on code paths that actually execute, so it is only as good as your tests’ concurrency. It cannot be combined with AddressSanitizer in the same build, so CI usually has one job for each. And every translation unit, ideally including dependencies, must be built with the flag, or you get false reports around uninstrumented code. MSVC has no ThreadSanitizer; on Windows, run the Linux build under WSL or use Clang.

When I start on a codebase with intermittent crashes, getting the test suite to run under TSan is usually the most productive first day of work: the reports are concrete, and the races they find are often the ones behind the “cannot reproduce” tickets.


Ten common bugs

1. Unsynchronized shared counter → atomic or mutex, shown above.

2. Concurrent vector mutation → protect all writers, and readers too if any writer exists.

std::vector<int> results;
std::mutex results_mtx;
void produce(int v) {
    std::lock_guard lock(results_mtx);
    results.push_back(v);  // push_back may reallocate: unsafe without the lock
}

Two unsynchronized push_back calls can both decide to reallocate, and one thread writes into a buffer the other just freed. The symptom is typically heap corruption detected much later, for example glibc aborting with malloc(): corrupted top size.

3. False sharing → separate hot atomics to different cache lines (alignas(64) patterns where justified).

struct alignas(64) PaddedCounter { std::atomic<long> value{0}; };
PaddedCounter per_thread[8];  // each counter on its own cache line

This is a performance bug, not a correctness bug: per-thread counters packed next to each other share a cache line, so every increment invalidates the other cores’ copy. C++17 offers std::hardware_destructive_interference_size, though not every standard library defines it; 64 bytes is the common value on x86.

4. Broken double-checked locking → std::call_once or static locals (since C++11).

Widget& instance() {
    static Widget w;  // initialized exactly once, thread-safe since C++11
    return w;
}

The hand-written version checks a plain pointer outside the lock, and another thread can see the pointer set before the object it points to is fully constructed.

5. Condition variable without predicate loop → handle spurious wakeups with wait(lock, pred).

std::unique_lock lock(mtx);
cv.wait(lock, [] { return !queue.empty(); });  // re-checks after every wakeup

wait can return without a notification, and a notification sent before the waiter started waiting is lost. The predicate handles both. condition_variable::wait needs a std::unique_lock because it unlocks and relocks the mutex internally, which lock_guard cannot do.

6. Reading shared state while another thread writes without synchronization. A bool stop = false; flag polled in a loop is the classic case: the compiler may hoist the load out of the loop and the thread never stops. Use std::atomic<bool>. volatile is not a substitute in C++; it does not create any ordering between threads.

7. Concurrent unique_ptr / shared_ptr reassignment. shared_ptr’s reference count is atomic, but the shared_ptr object itself is not: one thread assigning global_ptr = make_shared<T>() while another copies global_ptr is a data race. Protect it with a mutex or use C++20’s std::atomic<std::shared_ptr<T>>.

8. Iterator invalidation across threads. One thread iterates a container while another inserts or erases. Even if the writer holds a lock, the reader must hold it for the whole iteration, not just while fetching each element.

9. Non-thread-safe singleton patterns → call_once / Meyers singleton, as in bug 4.

10. Too-small critical sections split across logically atomic updates.

// ❌ each step is locked, the sequence is not
{ std::lock_guard l(m); if (cache.count(key)) return cache[key]; }
auto value = compute(key);
{ std::lock_guard l(m); cache[key] = value; }  // two threads may compute and insert

Whether this is acceptable depends on the invariant. Computing a value twice is often fine for a cache; debiting an account twice is not (see the withdraw example in section 1).

A last crash that looks like a race but is not one: destroying a std::thread that is still joinable calls std::terminate (see the FAQ below).


Patterns: thread pool sketch & shared_mutex

A work queue with mutex + condition_variable is the standard producer/consumer pattern: hold the lock only to take/push tasks; execute work outside the lock. Running the task while holding the queue lock serializes all workers, and if the task itself pushes another task onto the queue, it deadlocks on a non-recursive mutex.

std::shared_mutex: many concurrent readers or one writer—ideal for read-heavy caches if invariants are simple. Readers take std::shared_lock, writers std::unique_lock. It is not automatically faster than a plain mutex: acquiring a shared lock still writes to a shared counter, so for very short critical sections a std::mutex can win. Measure before switching.


Summary

Checklist

  • Every shared mutable object has a clear synchronization policy?
  • Reads synchronized when writes may occur?
  • Lock ordering documented to avoid deadlock?
  • Condition variables use predicates?
  • TSan runs on CI for threaded tests?

Tool choice

ScenarioTool
Simple counters/flagsatomic
Multi-field invariantsmutex
Read-mostly mapsshared_mutex (careful)
One-time initstd::call_once or a function-local static
Thread lifetimestd::jthread (C++20)

Rules

  1. No unsynchronized data races on non-atomic objects.
  2. Prefer scoped_lock for multiple mutexes.
  3. TSan on threaded test suites.
  4. Keep critical sections short, but large enough to cover the whole invariant; hold the lock for the full iteration when iterating shared containers.
  5. Document lock ordering for reviewers.

Concurrent bugs are hard to reproduce; ThreadSanitizer turns many races into actionable reports. Treat an intermittent failure in threaded code as a race until proven otherwise, and pair clear ownership of shared state with mutex/atomic discipline.


Frequently Asked Questions (FAQ)

Q. Why does my program abort when a std::thread object goes out of scope?

A. Destroying a std::thread that is still joinable calls std::terminate, so you must call join() or detach() on every path, including when an exception is thrown between creating the thread and joining it. That crash is easy to mistake for a race-related one. C++20’s std::jthread joins automatically in its destructor and also supports cooperative cancellation through std::stop_token.

Q. Is volatile enough for a flag shared between threads?

A. No. In C++ volatile only prevents the compiler from optimizing away accesses; it provides no atomicity and no ordering between threads, so accessing a volatile bool from two threads is still a data race. Use std::atomic<bool>.