Fixing Data Races with std::mutex: lock_guard, unique_lock, scoped_lock and Deadlock Avoidance

Introduction: Why did the counter break?

A typical version of this bug: several worker threads increment one shared order counter during a busy period, and after the batch the in-memory total no longer matches the number of rows in the database. Nothing crashed, no exception was logged, and the error only appears under load.

Cause: plain int counter + counter++ from many threads → lost updates (data race). ++counter is a load, an add and a store; two threads that load the same old value both write back old + 1, and one increment disappears. Fix: std::atomic for a single counter, or std::mutex when you update several fields consistently.

Mutex enforces mutual exclusion: only one thread runs the critical section at a time.

sequenceDiagram
  participant T1 as Thread 1
  participant M as mutex
  participant T2 as Thread 2
  T1->>M: lock()
  M-->>T1: acquired
  T2->>M: lock()
  Note over T2: blocked
  T1->>M: critical section
  T1->>M: unlock()
  M-->>T2: acquired
  T2->>M: critical section

Race condition and data race

Race condition: the outcome depends on scheduling order.

Data race (C++): concurrent unsynchronized access to the same memory where at least one access is a write → undefined behavior.

The two overlap but are not the same. A data race is a precise language rule, and tools such as ThreadSanitizer can detect it. A race condition is a logic error: every access can be perfectly synchronized and the program can still be wrong, because the sequence of operations was not protected as a whole. A mutex fixes both only if its critical section covers the whole sequence that must appear atomic, which is the theme of most bugs below.


Bug scenarios

  • Stock / check-then-act: guard both check and update with one lock
  • Log buffer: string += from many threads → protect with mutex
  • Work queue: empty then front/pop must be one atomic operation → mutex + often condition_variable
  • Hit/miss counters: update related stats under one lock for consistent snapshots
  • Config map: concurrent read/write on std::map → shared_mutex or external mutex

The first scenario is the one that survives code review most often:

// ❌ each call is locked, the sequence is not
bool reserve(Inventory& inv, int qty) {
    if (inv.available() >= qty) {   // lock, read, unlock
        inv.subtract(qty);          // lock, write, unlock
        return true;
    }
    return false;
}

// ✅ one lock around check and act
bool Inventory::reserve(int qty) {
    std::scoped_lock lock(mtx_);
    if (stock_ < qty) return false;
    stock_ -= qty;
    return true;
}

In the first version, two threads can both see available() == 5, both pass the check for 5 units, and the stock goes negative. Every individual method is “thread-safe”, which is exactly why the bug is hard to see. The fix moves the whole decision inside the class, under one lock. The same shape applies to the work queue: if (!q.empty()) { auto x = q.front(); q.pop(); } is broken even if empty, front and pop each lock, because another consumer can take the last item between the calls. A thread-safe queue therefore offers a single try_pop(T& out) instead.


std::mutex — basic usage

// Paste: g++ -std=c++17 -pthread -o mutex_safe mutex_safe.cpp && ./mutex_safe
#include <mutex>
#include <thread>
#include <iostream>
int counter = 0;
std::mutex counter_mutex;
void safeIncrement() {
    for (int i = 0; i < 100000; ++i) {
        std::lock_guard<std::mutex> lock(counter_mutex);
        ++counter;
    }
}
int main() {
    std::thread t1(safeIncrement);
    std::thread t2(safeIncrement);
    t1.join();
    t2.join();
    std::cout << "counter=" << counter << "\n";  // always 200000
    return 0;
}

Remove the lock_guard line and the result is usually below 200,000 and different on every run. With it, the result is always exactly 200,000. The mutex does two things here: it makes the read-modify-write exclusive, and it guarantees that each thread sees the other’s writes (unlocking synchronizes with the next lock), which is why no volatile or atomics are needed around counter.

Prefer lock_guard / unique_lock over raw lock()/unlock() so exceptions can’t leave the mutex locked. A manual unlock() at the end of a function is skipped by every early return and every exception, and the next thread that calls lock() then waits forever.

try_lock: a non-blocking attempt that returns false immediately if another thread holds the mutex. It is useful for “do this work if the resource is free, otherwise skip it” and for some deadlock-avoidance algorithms, but a loop that spins on try_lock() just burns CPU; if you need to wait, call lock(). To combine it with RAII, use std::unique_lock<std::mutex> lk(m, std::try_to_lock); if (lk.owns_lock()) { ... }.


lock_guard and unique_lock

RAII: lock in constructor, unlock in destructor.

unique_lock: can unlock()/lock() mid-scope—needed for condition_variable.

std::mutex m;
std::condition_variable cv;
std::queue<Job> jobs;

Job take() {
    std::unique_lock lock(m);
    cv.wait(lock, [] { return !jobs.empty(); });  // unlocks while waiting
    Job j = std::move(jobs.front());
    jobs.pop();
    return j;                                      // unlocks on return
}

condition_variable::wait must release the mutex while the thread sleeps, otherwise the producer could never lock it to push a job, and it must reacquire the mutex before returning. That is why it takes a unique_lock, which can be unlocked and relocked, and not a lock_guard, which cannot. The predicate form handles spurious wakeups and notifications that arrive before the consumer starts waiting. unique_lock also supports std::defer_lock (construct without locking), std::try_to_lock, and moving the lock out of a function. That flexibility costs a flag telling it whether it currently owns the mutex, so for plain scopes the simpler types are the default.

scoped_lock (C++17): lock multiple mutexes deadlock-free:

std::scoped_lock lock(m1, m2);

With a single mutex, scoped_lock behaves like lock_guard, so many codebases simply use scoped_lock everywhere in C++17 and later. One trap applies to all three: forgetting the variable name. std::lock_guard<std::mutex>{m}; creates a temporary that locks and immediately unlocks within that one statement, leaving the rest of the scope unprotected. std::scoped_lock(m); is worse: C++ parses it as the declaration of a new variable named m of type std::scoped_lock<>, which locks nothing and compiles without complaint. Always name the lock object (std::scoped_lock lock(m);).


Avoiding deadlock

Cause: two threads lock A then B vs B then A.

void transfer(Account& from, Account& to, int amount) {
    std::lock_guard a(from.m);   // thread 1: locks X, thread 2: locks Y
    std::lock_guard b(to.m);     // thread 1 waits for Y, thread 2 waits for X
    from.balance -= amount;
    to.balance += amount;
}
// thread 1: transfer(X, Y, 10);  thread 2: transfer(Y, X, 5);  → may hang forever

Neither thread made a mistake on its own: each locks “from” then “to”. The deadlock comes from the combination, and only when the timing lines up, which is why it tends to appear under production load and not in tests.

Fixes:

  1. Global lock order (always A then B): for example, always lock the account with the lower ID first. This works for any number of mutexes, but it is a convention that every code path must follow.
  2. std::lock(lock1, lock2) with defer_lock + unique_lock: std::lock acquires several mutexes using a deadlock-avoidance algorithm (lock one, try the others, back off and retry).
  3. std::scoped_lock(m1, m2) (recommended on C++17+): the same algorithm in one RAII object. Replacing the two lock_guard lines above with std::scoped_lock lock(from.m, to.m); removes the deadlock.
  4. Minimize critical section—no I/O inside locks, and never call unknown code (callbacks, virtual functions supplied by users) while holding a lock, because you cannot know which locks it takes.

When a deadlock does happen, it is one of the easier concurrency bugs to diagnose: the process is stuck, so attach a debugger (gdb -p <pid>, then thread apply all bt) or open the Parallel Stacks window in Visual Studio, and each blocked thread shows exactly which lock it is waiting for.


Common mistakes

  • return before unlock with manual lock/unlock → use RAII
  • Mutex far from data → easy to access data unlocked → encapsulate
  • Double-lock same std::mutex in one thread → undefined behavior, usually a self-deadlock (the mutex is non-recursive)
  • Callbacks under lock → can re-enter or block on other locks → call callbacks outside the lock
  • Returning references to protected data: a getter that returns const std::vector<T>& from inside a locked method gives the caller access after the lock is released. Return a copy, or take a function to run while the lock is held.
  • Locking a copy: a std::mutex cannot be copied, but a mutex passed around by pointer or captured by value in a wrapper can end up being a different mutex than the one other threads use. Two threads locking two different mutexes provide no protection at all.

The double-lock case usually arises through a call chain: public void add() locks, then calls public size(), which locks again. std::recursive_mutex makes that “work”, but it hides the design problem and makes it hard to reason about which invariants hold at each point. The cleaner fix is a private size_unlocked() helper that documents “caller must hold the lock”.


Performance

ApproachWhen
mutexMultiple variables, invariants across fields
atomicSingle counters/flags, no invariants with other vars
Lock-freeExpert-only; hard to get right

Keep lock duration minimal; don’t hold mutexes across I/O.

An uncontended lock and unlock is cheap on mainstream platforms (on Linux, std::mutex sits on a futex and only enters the kernel when a thread actually has to wait). The cost comes from contention: when many threads want the same lock, they queue, get descheduled and woken up, and the program effectively runs one thread at a time. The counter example above locks 200,000 times; under contention it can be slower than a single thread. The fix is to lock less often, not to find a faster mutex: count in a local variable and add the total under the lock once, or shard the data so different threads use different mutexes.

std::shared_mutex allows many readers or one writer. It helps when reads are frequent and relatively long; for very short critical sections its extra bookkeeping can make it slower than a plain std::mutex, so measure before switching.


Keeping the mutex inside the class it protects

Bundle data + mutex in one class and expose only thread-safe methods:

class Stats {
public:
    void record(bool hit) {
        std::scoped_lock lock(mtx_);
        (hit ? hits_ : misses_)++;
    }
    std::pair<long, long> snapshot() const {
        std::scoped_lock lock(mtx_);
        return {hits_, misses_};          // copy out under the lock
    }
private:
    mutable std::mutex mtx_;              // mutable: lock in const methods
    long hits_ = 0, misses_ = 0;
};

Because callers never see the mutex, they cannot forget to lock it, and snapshot() returns a consistent pair: hits and misses from the same moment, which two separate atomics could not guarantee. The mutex is mutable so that const observers can still lock it.

  • Thread-safe queue wrapper: all methods lock internally, and combined operations (try_pop) replace check-then-act sequences.
  • shared_mutex: std::shared_lock for readers, std::unique_lock for writers, for read-mostly caches and configuration.
  • Snapshot publish: for data that is read constantly and replaced rarely, readers load a std::shared_ptr<const Config> and writers publish a new one. C++20 provides std::atomic<std::shared_ptr<T>> for this; the older free functions std::atomic_load/std::atomic_store on shared_ptr are deprecated in C++20.

Run the test suite under ThreadSanitizer (g++ -fsanitize=thread -g) to catch any access that bypasses the mutex.


References