C++ std::atomic: Atomic Operations, Memory Order, and compare_exchange
Key takeaways
std::atomic makes a single read-modify-write indivisible, and its memory_order argument decides what other memory becomes visible along with it. Covers the atomic operations, a memory_order table with the bugs each wrong choice causes, compare_exchange loops, and where a mutex is simpler.
What is an atomic operation?
An atomic operation is indivisible: no other thread can observe it half-done. In C++, std::atomic<T> provides such operations, and using it for shared variables that are written by one thread and read by another turns what would be a data race (undefined behavior) into well-defined behavior.
#include <atomic>
std::atomic<int> counter{0};
counter++; // one indivisible read-modify-write
int counter2 = 0;
counter2++; // load, add, store: three steps, racy if shared
Why the second one breaks:
int counter = 0;
void increment() {
counter++; // 1. load, 2. add, 3. store
}
// Thread 1: load(0) -> add(1) -> store(1)
// Thread 2: load(0) -> add(1) -> store(1)
// Result: 1 (expected 2)
With std::atomic<int>, counter++ compiles to a single hardware read-modify-write (for example lock xadd on x86), so both increments are kept.
flowchart TD
A[Thread 1: counter++] --> B{atomic?}
B -->|Yes| C[single hardware RMW instruction]
B -->|No| D[1. load]
D --> E[2. add]
E --> F[3. store]
C --> G[complete]
F --> H{Thread 2 in between?}
H -->|Yes| I[lost update]
H -->|No| G
Atomicity is only half of the story. The other half is ordering: when thread A writes some ordinary data and then sets an atomic flag, does thread B, after seeing the flag, also see the data? That is what the memory_order argument controls, covered below.
Atomic vs mutex
std::atomic | std::mutex | |
|---|---|---|
| Protects | One variable, one operation at a time | Any number of variables, any code |
| Cost (uncontended) | One atomic instruction | Atomic lock + unlock, possibly a syscall under contention |
| Blocking | Never sleeps (lock-free types) | Waiting threads sleep |
| Deadlock | Not possible | Possible with multiple locks |
| Difficulty | Easy for counters/flags, very hard for data structures | Easy to reason about |
// atomic: a single counter
std::atomic<int> counter{0};
counter++;
// mutex: an invariant spanning a container
std::mutex mtx;
std::map<int, int> data;
{
std::lock_guard lock(mtx);
data[key] = value;
}
A good rule of thumb: if the correctness argument involves two variables at once (“the size matches the contents”, “the head points to a node whose next is valid”), reach for a mutex first. Atomics compose badly; two individually atomic operations are not atomic together.
The operations
#include <atomic>
std::atomic<int> x{0};
x.store(10); // write
int v = x.load(); // read
int old = x.exchange(20); // write, return previous value
x.fetch_add(5); // returns the old value
x.fetch_sub(3);
x.fetch_and(0xFF); // integral types only
x.fetch_or(0x10);
x++; x--; x += 2; // operators, all seq_cst
int expected = 24;
x.compare_exchange_strong(expected, 30); // CAS, see below
std::atomic<T> works for any trivially copyable T: integers, pointers, bool, small structs, and since C++20 float/double also get fetch_add. It does not work for std::string or std::vector (they are not trivially copyable, so std::atomic<std::string> fails to compile). For those, use a mutex, or std::atomic<std::shared_ptr<T>> (C++20) to swap whole immutable snapshots.
A Counter and a Done Flag
Counter
#include <atomic>
#include <iostream>
#include <thread>
#include <vector>
std::atomic<int> counter{0};
void increment() {
for (int i = 0; i < 1000; i++) {
counter.fetch_add(1, std::memory_order_relaxed);
}
}
int main() {
std::vector<std::thread> threads;
for (int i = 0; i < 10; i++) threads.emplace_back(increment);
for (auto& t : threads) t.join();
std::cout << counter << '\n'; // always 10000
}
relaxed is correct here because nothing else is published through the counter, and join() already synchronizes the final read.
A “done” flag that publishes a result
#include <atomic>
#include <chrono>
#include <iostream>
#include <thread>
int result = 0; // plain int, published via the flag
std::atomic<bool> done{false};
void worker() {
result = 42;
done.store(true, std::memory_order_release);
}
void monitor() {
while (!done.load(std::memory_order_acquire)) {
std::this_thread::sleep_for(std::chrono::milliseconds(10));
}
std::cout << result << '\n'; // guaranteed 42
}
The release store and the acquire load that reads true form a synchronizes-with edge: everything the worker wrote before the store is visible to the monitor after the load. With relaxed on either side, reading result would be a data race. For waits longer than a few milliseconds, a condition_variable or C++20 done.wait(false) avoids polling.
Memory order
Compilers and CPUs reorder memory accesses as long as a single thread can’t tell. Other threads can tell. The memory_order argument on each atomic operation limits that reordering.
| Order | Applies to | Guarantee | Typical use |
|---|---|---|---|
relaxed | any | Atomicity only; no ordering with other memory | Statistics counters, reference count increments |
acquire | loads | Reads/writes after the load can’t move before it | Taking a lock, reading a “ready” flag |
release | stores | Reads/writes before the store can’t move after it | Releasing a lock, setting a “ready” flag |
acq_rel | read-modify-write | Both of the above | CAS in lock-free structures, reference count decrement |
seq_cst | any (default) | acquire/release, plus one global order of all seq_cst operations that every thread agrees on | When in doubt; algorithms that need a total order |
(memory_order_consume also exists, but every major compiler treats it as acquire, and its use is discouraged.)
relaxed: atomic, but no ordering
std::atomic<int> x{0}, y{0};
void thread1() {
x.store(1, std::memory_order_relaxed);
y.store(1, std::memory_order_relaxed);
}
void thread2() {
while (y.load(std::memory_order_relaxed) == 0) {}
int r = x.load(std::memory_order_relaxed); // may be 0
}
Seeing y == 1 says nothing about x. On x86 this happens to “work” because the hardware keeps stores in order, which is exactly why relaxed-order bugs survive testing on developer laptops and then appear on ARM servers or phones. The compiler is also allowed to reorder the two stores, on any architecture.
acquire / release: publishing data
std::atomic<int> data{0};
std::atomic<bool> ready{false};
void producer() {
data.store(42, std::memory_order_relaxed);
ready.store(true, std::memory_order_release);
}
void consumer() {
while (!ready.load(std::memory_order_acquire)) {}
int v = data.load(std::memory_order_relaxed); // guaranteed 42
}
The pairing is what matters: a release store only helps a thread that does an acquire load of the same atomic and reads the value written by that store. Using release on the producer and relaxed on the consumer gives you nothing.
seq_cst: when acquire/release is not enough
The classic case where acquire/release is too weak is two threads that each write one flag and then read the other’s:
std::atomic<bool> a{false}, b{false};
int r1, r2;
void t1() { a.store(true); r1 = b.load(); }
void t2() { b.store(true); r2 = a.load(); }
With the default seq_cst, at least one of r1, r2 is true: there is a single order of all four operations and one of the stores must come first. With release stores and acquire loads, r1 == r2 == false is allowed, because each CPU’s store can sit in its store buffer while its load runs. This is the pattern behind Dekker-style mutual exclusion and some “is anyone else in here?” checks; if your algorithm looks like this, keep seq_cst.
On x86, seq_cst loads cost the same as acquire loads, and only seq_cst stores are more expensive (an xchg or mfence). On ARM the difference is smaller than it used to be since ARMv8 added dedicated acquire/release instructions. Start with the default and weaken only with a clear reason and a test on weakly ordered hardware.
compare_exchange (CAS)
compare_exchange_*(expected, desired) does this atomically: if the current value equals expected, replace it with desired and return true; otherwise write the current value into expected and return false.
std::atomic<int> value{0};
int expected = 0;
if (value.compare_exchange_strong(expected, 10)) {
// value was 0, now 10
} else {
// value was not 0; expected now holds what it actually was
}
The update of expected on failure is what makes the standard retry loop short. For example, an atomic “maximum” (there is no fetch_max in the standard):
void update_max(std::atomic<int>& maximum, int candidate) {
int current = maximum.load(std::memory_order_relaxed);
while (current < candidate &&
!maximum.compare_exchange_weak(current, candidate,
std::memory_order_relaxed)) {
// current was refreshed by the failed CAS; loop re-checks it
}
}
weak vs strong. compare_exchange_weak is allowed to fail even when the value matches (a spurious failure). On ARM and other load-linked/store-conditional architectures, that lets it compile to a single LL/SC pair, while strong needs an inner loop. Inside a retry loop use weak; for a single attempt where a false failure would be wrong, use strong.
Two orderings. CAS accepts separate orders for success and failure: compare_exchange_weak(expected, desired, success_order, failure_order). The failure order applies to the plain load that happens when the comparison fails, so it can’t be release or acq_rel, and it’s commonly relaxed.
Example: lock-free stack push
#include <atomic>
template <typename T>
class LockFreeStack {
struct Node {
T data;
Node* next;
};
std::atomic<Node*> head_{nullptr};
public:
void push(T value) {
Node* node = new Node{std::move(value), head_.load(std::memory_order_relaxed)};
while (!head_.compare_exchange_weak(node->next, node,
std::memory_order_release,
std::memory_order_relaxed)) {
// node->next now holds the current head; try again
}
}
};
The release on success publishes the node’s contents to whichever thread later acquires head_. Passing node->next as expected means a failed CAS automatically re-links the node to the new head.
pop is where lock-free stacks get genuinely hard. The textbook version
Node* old = head_.load(std::memory_order_acquire);
while (old && !head_.compare_exchange_weak(old, old->next,
std::memory_order_acquire,
std::memory_order_relaxed)) {}
// ... use old->data, then delete old
has two bugs once more than one thread pops. First, old->next is read from a node that another thread may already have popped and deleted: a use-after-free. Second, the ABA problem: thread 1 reads head A with next B; thread 2 pops A, pops B, pushes A back; thread 1’s CAS sees A again, succeeds, and installs B, which has been freed. Fixes are a tagged pointer (pointer plus a counter compared together, which needs a 16-byte CAS: cmpxchg16b on x86-64, -mcx16 on GCC), or a safe memory reclamation scheme such as hazard pointers or epoch-based reclamation. In practice, use a tested library (Boost.Lockfree, folly, moodycamel’s queue) rather than writing this yourself.
Non-Atomic Compound Expressions, volatile, and False Sharing
Compound expressions are not atomic
std::atomic<int> x{0};
x = x + 1; // ❌ atomic load, then atomic store: another thread can slip in between
x++; // ✅
x.fetch_add(1); // ✅
Each access in x = x + 1 is atomic on its own, but the pair is not. The same applies to if (x.load() == 0) x.store(1);, which should be a CAS.
Relaxed flag publishing plain data
std::atomic<bool> ready{false};
int data = 0;
void producer() {
data = 42;
ready.store(true, std::memory_order_relaxed); // ❌
}
void consumer() {
while (!ready.load(std::memory_order_relaxed)) {}
std::cout << data; // data race: undefined behavior, may print 0
}
Change the store to release and the load to acquire. ThreadSanitizer (-fsanitize=thread) reports this kind of race reliably; it’s the fastest way to check a new synchronization scheme.
Large types are not lock-free
std::atomic<int> x;
std::cout << x.is_lock_free(); // 1 on mainstream platforms
static_assert(std::atomic<int>::is_always_lock_free); // C++17, compile-time
struct Large { int data[100]; };
std::atomic<Large> large; // compiles, but uses a hidden lock
std::atomic<Large> falls back to a lock internally. With GCC this also means linking with -latomic, otherwise you get undefined references to __atomic_load. If you need lock-free behavior, check is_always_lock_free at compile time rather than trusting that it is.
volatile is not atomic
volatile int counter = 0;
counter++; // ❌ still a racy load/add/store
std::atomic<int> counter2{0};
counter2++; // ✅
volatile only prevents the compiler from optimizing away accesses; it’s meant for memory-mapped hardware registers. It gives neither atomicity nor ordering between threads (MSVC’s old /volatile:ms extension is an exception, but don’t rely on it).
False sharing
Two atomics that are updated by different threads but sit on the same 64-byte cache line slow each other down, because each write invalidates the line for the other core. Per-thread counters that are summed later should be padded apart:
struct alignas(64) PaddedCounter {
std::atomic<long> value{0};
};
PaddedCounter per_thread[8];
C++17 provides std::hardware_destructive_interference_size for the constant, though GCC warns about using it in headers because its value can differ between compiler flags.
Spinlock, Lazy Initialization, and Reference Counting
Spinlock
#include <atomic>
#include <thread>
class SpinLock {
std::atomic<bool> locked_{false};
public:
void lock() {
for (;;) {
if (!locked_.exchange(true, std::memory_order_acquire)) return;
while (locked_.load(std::memory_order_relaxed)) {
std::this_thread::yield(); // wait without hammering the cache line
}
}
}
void unlock() { locked_.store(false, std::memory_order_release); }
};
This is the “test and test-and-set” form: spin on a plain load, and only attempt the exchange when the lock looks free. A loop that calls exchange continuously forces the cache line to move between cores on every iteration. Spinlocks only make sense for very short critical sections; if the holder can be preempted, waiting threads burn CPU for nothing, and std::mutex, which puts waiters to sleep, is usually the better choice.
Lazy initialization
Double-checked locking with an atomic pointer is correct when written with acquire/release:
#include <atomic>
#include <mutex>
class Singleton {
static std::atomic<Singleton*> instance_;
static std::mutex mutex_;
Singleton() = default;
public:
static Singleton* getInstance() {
Singleton* tmp = instance_.load(std::memory_order_acquire);
if (tmp == nullptr) {
std::lock_guard lock(mutex_);
tmp = instance_.load(std::memory_order_relaxed);
if (tmp == nullptr) {
tmp = new Singleton();
instance_.store(tmp, std::memory_order_release);
}
}
return tmp;
}
};
std::atomic<Singleton*> Singleton::instance_{nullptr};
std::mutex Singleton::mutex_;
But since C++11 you almost never need it. A function-local static is initialized thread-safely by the compiler, with the same fast path:
Singleton& Singleton::get() {
static Singleton instance; // initialized exactly once, thread-safe
return instance;
}
For non-static cases, std::call_once does the same job.
Reference counting
std::shared_ptr implementations use essentially this scheme:
struct RefCounted {
std::atomic<int> refs{1};
void add_ref() { refs.fetch_add(1, std::memory_order_relaxed); }
void release() {
if (refs.fetch_sub(1, std::memory_order_acq_rel) == 1) {
delete this;
}
}
};
Increments can be relaxed because a thread can only add a reference to an object it already holds. The decrement needs acq_rel so that all writes by other owners happen-before the delete in whichever thread drops the last reference.
FAQ
When should I use atomic?
For single counters and flags, reference counts, and as the building block inside lock-free structures written by people who test them on weak-memory hardware. For anything else, a mutex is easier to get right.
How do I choose a memory_order?
Use the default (seq_cst) unless you can name the happens-before edge you need. Use relaxed for counters that publish nothing, and release/acquire pairs for “write data, then set flag” hand-offs. Test on ARM, or at least run ThreadSanitizer, before trusting weaker orders.
Is lock-free always faster?
No. Lock-free means some thread always makes progress, not that each thread is fast. Under heavy contention a CAS loop can retry many times, and a mutex that puts waiters to sleep may give better throughput.
How do I debug atomics?
ThreadSanitizer finds data races on non-atomic data and many ordering mistakes. Logging tends to change timing enough to hide the bug. Model checkers such as CDSChecker or Relacy can exhaustively explore interleavings for small lock-free algorithms.