Fixing Data Races with std::mutex: lock_guard, unique_lock, scoped_lock and Deadlock Avoidance
Introduction: Why did the counter break?
A typical version of this bug: several worker threads increment one shared order counter during a busy period, and after the batch the in-memory total no longer matches the number of rows in the database. Nothing crashed, no exception was logged, and the error only appears under load.
Cause: plain int counter + counter++ from many threads → lost updates (data race). ++counter is a load, an add and a store; two threads that load the same old value both write back old + 1, and one increment disappears. Fix: std::atomic
Mutex enforces mutual exclusion: only one thread runs the critical section at a time.
sequenceDiagram participant T1 as Thread 1 participant M as mutex participant T2 as Thread 2 T1->>M: lock() M-->>T1: acquired T2->>M: lock() Note over T2: blocked T1->>M: critical section T1->>M: unlock() M-->>T2: acquired T2->>M: critical section
Race condition and data race
Race condition: the outcome depends on scheduling order.
Data race (C++): concurrent unsynchronized access to the same memory where at least one access is a write → undefined behavior.
The two overlap but are not the same. A data race is a precise language rule, and tools such as ThreadSanitizer can detect it. A race condition is a logic error: every access can be perfectly synchronized and the program can still be wrong, because the sequence of operations was not protected as a whole. A mutex fixes both only if its critical section covers the whole sequence that must appear atomic, which is the theme of most bugs below.
Bug scenarios
- Stock / check-then-act: guard both check and update with one lock
- Log buffer:
string +=from many threads → protect with mutex - Work queue:
emptythenfront/popmust be one atomic operation → mutex + oftencondition_variable - Hit/miss counters: update related stats under one lock for consistent snapshots
- Config map: concurrent read/write on
std::map→shared_mutexor external mutex
The first scenario is the one that survives code review most often:
// ❌ each call is locked, the sequence is not
bool reserve(Inventory& inv, int qty) {
if (inv.available() >= qty) { // lock, read, unlock
inv.subtract(qty); // lock, write, unlock
return true;
}
return false;
}
// ✅ one lock around check and act
bool Inventory::reserve(int qty) {
std::scoped_lock lock(mtx_);
if (stock_ < qty) return false;
stock_ -= qty;
return true;
}
In the first version, two threads can both see available() == 5, both pass the check for 5 units, and the stock goes negative. Every individual method is “thread-safe”, which is exactly why the bug is hard to see. The fix moves the whole decision inside the class, under one lock. The same shape applies to the work queue: if (!q.empty()) { auto x = q.front(); q.pop(); } is broken even if empty, front and pop each lock, because another consumer can take the last item between the calls. A thread-safe queue therefore offers a single try_pop(T& out) instead.
std::mutex — basic usage
// Paste: g++ -std=c++17 -pthread -o mutex_safe mutex_safe.cpp && ./mutex_safe
#include <mutex>
#include <thread>
#include <iostream>
int counter = 0;
std::mutex counter_mutex;
void safeIncrement() {
for (int i = 0; i < 100000; ++i) {
std::lock_guard<std::mutex> lock(counter_mutex);
++counter;
}
}
int main() {
std::thread t1(safeIncrement);
std::thread t2(safeIncrement);
t1.join();
t2.join();
std::cout << "counter=" << counter << "\n"; // always 200000
return 0;
}
Remove the lock_guard line and the result is usually below 200,000 and different on every run. With it, the result is always exactly 200,000. The mutex does two things here: it makes the read-modify-write exclusive, and it guarantees that each thread sees the other’s writes (unlocking synchronizes with the next lock), which is why no volatile or atomics are needed around counter.
Prefer lock_guard / unique_lock over raw lock()/unlock() so exceptions can’t leave the mutex locked. A manual unlock() at the end of a function is skipped by every early return and every exception, and the next thread that calls lock() then waits forever.
try_lock: a non-blocking attempt that returns false immediately if another thread holds the mutex. It is useful for “do this work if the resource is free, otherwise skip it” and for some deadlock-avoidance algorithms, but a loop that spins on try_lock() just burns CPU; if you need to wait, call lock(). To combine it with RAII, use std::unique_lock<std::mutex> lk(m, std::try_to_lock); if (lk.owns_lock()) { ... }.
lock_guard and unique_lock
RAII: lock in constructor, unlock in destructor.
unique_lock: can unlock()/lock() mid-scope—needed for condition_variable.
std::mutex m;
std::condition_variable cv;
std::queue<Job> jobs;
Job take() {
std::unique_lock lock(m);
cv.wait(lock, [] { return !jobs.empty(); }); // unlocks while waiting
Job j = std::move(jobs.front());
jobs.pop();
return j; // unlocks on return
}
condition_variable::wait must release the mutex while the thread sleeps, otherwise the producer could never lock it to push a job, and it must reacquire the mutex before returning. That is why it takes a unique_lock, which can be unlocked and relocked, and not a lock_guard, which cannot. The predicate form handles spurious wakeups and notifications that arrive before the consumer starts waiting. unique_lock also supports std::defer_lock (construct without locking), std::try_to_lock, and moving the lock out of a function. That flexibility costs a flag telling it whether it currently owns the mutex, so for plain scopes the simpler types are the default.
scoped_lock (C++17): lock multiple mutexes deadlock-free:
std::scoped_lock lock(m1, m2);
With a single mutex, scoped_lock behaves like lock_guard, so many codebases simply use scoped_lock everywhere in C++17 and later. One trap applies to all three: forgetting the variable name. std::lock_guard<std::mutex>{m}; creates a temporary that locks and immediately unlocks within that one statement, leaving the rest of the scope unprotected. std::scoped_lock(m); is worse: C++ parses it as the declaration of a new variable named m of type std::scoped_lock<>, which locks nothing and compiles without complaint. Always name the lock object (std::scoped_lock lock(m);).
Avoiding deadlock
Cause: two threads lock A then B vs B then A.
void transfer(Account& from, Account& to, int amount) {
std::lock_guard a(from.m); // thread 1: locks X, thread 2: locks Y
std::lock_guard b(to.m); // thread 1 waits for Y, thread 2 waits for X
from.balance -= amount;
to.balance += amount;
}
// thread 1: transfer(X, Y, 10); thread 2: transfer(Y, X, 5); → may hang forever
Neither thread made a mistake on its own: each locks “from” then “to”. The deadlock comes from the combination, and only when the timing lines up, which is why it tends to appear under production load and not in tests.
Fixes:
- Global lock order (always A then B): for example, always lock the account with the lower ID first. This works for any number of mutexes, but it is a convention that every code path must follow.
- std::lock(lock1, lock2) with
defer_lock+unique_lock:std::lockacquires several mutexes using a deadlock-avoidance algorithm (lock one, try the others, back off and retry). - std::scoped_lock(m1, m2) (recommended on C++17+): the same algorithm in one RAII object. Replacing the two
lock_guardlines above withstd::scoped_lock lock(from.m, to.m);removes the deadlock. - Minimize critical section—no I/O inside locks, and never call unknown code (callbacks, virtual functions supplied by users) while holding a lock, because you cannot know which locks it takes.
When a deadlock does happen, it is one of the easier concurrency bugs to diagnose: the process is stuck, so attach a debugger (gdb -p <pid>, then thread apply all bt) or open the Parallel Stacks window in Visual Studio, and each blocked thread shows exactly which lock it is waiting for.
Common mistakes
returnbeforeunlockwith manual lock/unlock → use RAII- Mutex far from data → easy to access data unlocked → encapsulate
- Double-lock same
std::mutexin one thread → undefined behavior, usually a self-deadlock (the mutex is non-recursive) - Callbacks under lock → can re-enter or block on other locks → call callbacks outside the lock
- Returning references to protected data: a getter that returns
const std::vector<T>&from inside a locked method gives the caller access after the lock is released. Return a copy, or take a function to run while the lock is held. - Locking a copy: a
std::mutexcannot be copied, but a mutex passed around by pointer or captured by value in a wrapper can end up being a different mutex than the one other threads use. Two threads locking two different mutexes provide no protection at all.
The double-lock case usually arises through a call chain: public void add() locks, then calls public size(), which locks again. std::recursive_mutex makes that “work”, but it hides the design problem and makes it hard to reason about which invariants hold at each point. The cleaner fix is a private size_unlocked() helper that documents “caller must hold the lock”.
Performance
| Approach | When |
|---|---|
mutex | Multiple variables, invariants across fields |
atomic | Single counters/flags, no invariants with other vars |
| Lock-free | Expert-only; hard to get right |
Keep lock duration minimal; don’t hold mutexes across I/O.
An uncontended lock and unlock is cheap on mainstream platforms (on Linux, std::mutex sits on a futex and only enters the kernel when a thread actually has to wait). The cost comes from contention: when many threads want the same lock, they queue, get descheduled and woken up, and the program effectively runs one thread at a time. The counter example above locks 200,000 times; under contention it can be slower than a single thread. The fix is to lock less often, not to find a faster mutex: count in a local variable and add the total under the lock once, or shard the data so different threads use different mutexes.
std::shared_mutex allows many readers or one writer. It helps when reads are frequent and relatively long; for very short critical sections its extra bookkeeping can make it slower than a plain std::mutex, so measure before switching.
Keeping the mutex inside the class it protects
Bundle data + mutex in one class and expose only thread-safe methods:
class Stats {
public:
void record(bool hit) {
std::scoped_lock lock(mtx_);
(hit ? hits_ : misses_)++;
}
std::pair<long, long> snapshot() const {
std::scoped_lock lock(mtx_);
return {hits_, misses_}; // copy out under the lock
}
private:
mutable std::mutex mtx_; // mutable: lock in const methods
long hits_ = 0, misses_ = 0;
};
Because callers never see the mutex, they cannot forget to lock it, and snapshot() returns a consistent pair: hits and misses from the same moment, which two separate atomics could not guarantee. The mutex is mutable so that const observers can still lock it.
- Thread-safe queue wrapper: all methods lock internally, and combined operations (
try_pop) replace check-then-act sequences. - shared_mutex:
std::shared_lockfor readers,std::unique_lockfor writers, for read-mostly caches and configuration. - Snapshot publish: for data that is read constantly and replaced rarely, readers load a
std::shared_ptr<const Config>and writers publish a new one. C++20 providesstd::atomic<std::shared_ptr<T>>for this; the older free functionsstd::atomic_load/std::atomic_storeonshared_ptrare deprecated in C++20.
Run the test suite under ThreadSanitizer (g++ -fsanitize=thread -g) to catch any access that bypasses the mutex.