std::condition_variable: Waiting on Events Without Spurious Wakeup and Lost Notify Bugs

Key takeaways

A condition_variable carries no state of its own: it only lets a thread sleep until another thread changes state guarded by a mutex. Waiting with a predicate, modifying the state under the lock, and giving every waiter a shutdown condition prevent almost every bug people hit with it.

std::condition_variable lets a thread sleep until some other thread changes shared state, without burning a core polling for that change. The API is only a handful of functions (wait, wait_for, wait_until, notify_one, notify_all), but most of the bugs come from the rules the API leaves to you: which mutex protects which state, when the state may change, and how waiters find out it is time to stop.

This article goes through those rules in the order you usually run into them. All examples compile with g++ 10 (-std=c++17 -pthread; the stop_token example needs -std=c++20).

The three pieces: state, mutex, condition variable

The most important thing to understand is that a condition variable does not remember anything. A notification is not a message that waits in a mailbox. If nobody is waiting when notify_one() is called, it does nothing at all.

So every correct use has three parts:

  1. Shared state that expresses the condition (bool ready, a queue, a counter).
  2. A mutex that protects that state.
  3. The condition variable, which is only a way to sleep until the state might have changed.
#include <condition_variable>
#include <iostream>
#include <mutex>
#include <thread>

std::mutex mtx;
std::condition_variable cv;
bool ready = false;  // the shared state; the cv itself carries no data

void worker() {
    std::unique_lock<std::mutex> lock(mtx);
    cv.wait(lock, [] { return ready; });
    // lock is held again here, and ready is guaranteed to be true
    std::cout << "worker: ready observed\n";
}

int main() {
    std::thread t(worker);
    {
        std::lock_guard<std::mutex> lock(mtx);
        ready = true;
    }
    cv.notify_one();
    t.join();
}

Output:

worker: ready observed

wait(lock, pred) behaves like this loop:

while (!pred()) {
    cv.wait(lock);  // atomically unlock, sleep, and re-lock on wake-up
}

That is why it needs a std::unique_lock: it has to release the mutex while the thread sleeps and take it back before returning. std::lock_guard cannot be unlocked, so it cannot be passed in. Plain std::condition_variable also only accepts std::unique_lock<std::mutex>; for any other lock type you need std::condition_variable_any, covered below.

Note that the example does not need a sleep_for in main to “let the worker start first.” If the main thread sets ready before the worker reaches wait, the predicate is already true and the worker never sleeps. That property comes entirely from checking the predicate.

Spurious wakeups: why the predicate is not optional

The standard allows wait() to return even though nobody called notify. These spurious wakeups come from the underlying OS primitives (POSIX explicitly allows pthread_cond_wait to wake spuriously). There is also a closely related case that is not spurious at all: thread A is notified, but before it gets the mutex back, thread B takes the item it was woken for. By the time A runs, the condition is false again.

Both cases have the same fix: always re-check the condition after waking. The predicate overload does this for you, so the only form of wait I use in application code is wait(lock, pred). The overload without a predicate is easy to call and almost never correct on its own.

Lost notifications: modify the state under the lock

The predicate also protects against notify-before-wait, as long as the notifying side follows one rule: change the shared state while holding the same mutex the waiter uses.

This rule bites people who make the flag a std::atomic<bool> and conclude that the mutex is no longer needed on the write side:

std::atomic<bool> ready{false};

// Waiter                                  // Notifier (buggy)
std::unique_lock<std::mutex> lock(mtx);
// pred() reads ready == false ...
//                                         ready = true;      // no lock taken
//                                         cv.notify_one();   // nobody is sleeping yet
cv.wait(lock);  // ... and now sleeps, possibly forever

The waiter checked the predicate and was about to go to sleep. Between those two steps the notifier changed the flag and notified. Nobody was sleeping on the cv yet, so the notification disappeared. If the notifier had to lock mtx to set the flag, it could not have run in that window, because the waiter holds the mutex until wait has atomically released it and started sleeping.

I have seen this pattern survive testing for a long time because the window is tiny: it only shows up as an occasional hang under load, and adding logging usually changes the timing enough to hide it. When a waiter hangs “for no reason”, the first thing I check now is whether every write to the predicate’s state happens under the mutex.

The following program deliberately notifies before anyone waits. The notification is dropped, but the predicate check still sees ready == true and returns right away:

#include <chrono>
#include <condition_variable>
#include <iostream>
#include <mutex>
#include <thread>
using namespace std::chrono_literals;

std::mutex mtx;
std::condition_variable cv;
bool ready = false;

int main() {
    { std::lock_guard<std::mutex> lock(mtx); ready = true; }
    cv.notify_one();   // nobody is waiting yet: this notification is dropped
    std::thread t([] {
        std::unique_lock<std::mutex> lock(mtx);
        // cv.wait(lock);  // without a predicate this would block forever
        bool ok = cv.wait_for(lock, 1s, [] { return ready; });
        std::cout << "predicate wait returned " << std::boolalpha << ok << '\n';
    });
    t.join();
}
predicate wait returned true

Notify inside or outside the lock?

Both are correct as long as the state was modified under the lock. Notifying after unlocking avoids waking a thread only for it to block on a mutex that is still held (some implementations optimize this case, but you do not need to rely on it). Notifying while still holding the lock is the safer choice in one situation: when the waiter may destroy the condition variable as soon as it wakes, for example when the cv lives in a stack object that the waiter owns. Destroying a condition variable while any thread is still waiting on it is undefined behavior.

A bounded blocking queue: notify_one, notify_all and two condition variables

A work queue is the most common real use. This version has a capacity limit, so fast producers block instead of letting the queue grow without bound, and a close() operation so consumers can exit:

#include <condition_variable>
#include <iostream>
#include <mutex>
#include <optional>
#include <queue>
#include <thread>
#include <vector>

template <typename T>
class BlockingQueue {
public:
    explicit BlockingQueue(std::size_t capacity) : capacity_(capacity) {}

    // Returns false if the queue was closed before the item could be pushed.
    bool push(T value) {
        std::unique_lock<std::mutex> lock(mtx_);
        not_full_.wait(lock, [&] { return closed_ || q_.size() < capacity_; });
        if (closed_) return false;
        q_.push(std::move(value));
        lock.unlock();
        not_empty_.notify_one();
        return true;
    }

    // Returns std::nullopt once the queue is closed and drained.
    std::optional<T> pop() {
        std::unique_lock<std::mutex> lock(mtx_);
        not_empty_.wait(lock, [&] { return closed_ || !q_.empty(); });
        if (q_.empty()) return std::nullopt;  // closed and drained
        T value = std::move(q_.front());
        q_.pop();
        lock.unlock();
        not_full_.notify_one();
        return value;
    }

    void close() {
        {
            std::lock_guard<std::mutex> lock(mtx_);
            closed_ = true;
        }
        not_empty_.notify_all();
        not_full_.notify_all();
    }

private:
    std::mutex mtx_;
    std::condition_variable not_empty_;
    std::condition_variable not_full_;
    std::queue<T> q_;
    std::size_t capacity_;
    bool closed_ = false;
};

int main() {
    BlockingQueue<int> q(4);
    std::vector<std::thread> consumers;
    std::mutex out_mtx;
    long long total = 0;

    for (int id = 0; id < 3; ++id) {
        consumers.emplace_back([&] {
            long long local = 0;
            while (auto item = q.pop()) local += *item;
            std::lock_guard<std::mutex> lock(out_mtx);
            total += local;
        });
    }

    for (int i = 1; i <= 1000; ++i) q.push(i);
    q.close();

    for (auto& t : consumers) t.join();
    std::cout << "sum = " << total << '\n';
}
sum = 500500

Some design decisions in this queue are worth spelling out.

One item, one waiter: notify_one. Each push adds exactly one item, so waking one consumer is enough. Waking all of them would make every consumer take the mutex, see the queue empty again (except the one that won), and go back to sleep, which is wasted work under contention.

Two different conditions, two condition variables. Producers wait for “not full” and consumers wait for “not empty”. If both waited on the same cv, a notify_one from pop meant for a producer could wake another consumer instead. That consumer sees an empty queue and goes back to sleep, the producer is never woken, and the program stalls. This is a classic way to get a hang with no deadlock in the lock graph. If you really want a single cv for several different predicates, you have to use notify_all so the thread that can make progress is guaranteed to be among those woken.

Shutdown is part of the predicate. Both predicates include closed_. Without that, close() could notify all it wants and consumers blocked on an empty queue would re-check !q_.empty(), find it false, and sleep forever. notify_all is the right call here because every waiter needs to see the state change. Note that pop still drains remaining items after close, and returns nullopt only once the queue is actually empty.

The mistake I see most often in hand-written worker pools is the missing shutdown condition: workers wait on “queue not empty”, the destructor sets a stop flag and calls notify_all, and the program hangs at exit because the predicate never looks at stop. The other version of the same bug is join()ing workers that are still blocked in wait without notifying them at all.

Timeouts: wait_for and wait_until

Both timed functions take a predicate as well and return the predicate’s value when they return:

#include <chrono>
#include <condition_variable>
#include <iostream>
#include <mutex>
#include <thread>
using namespace std::chrono_literals;

std::mutex mtx;
std::condition_variable cv;
bool ready = false;

int main() {
    std::thread t([] {
        std::unique_lock<std::mutex> lock(mtx);
        auto deadline = std::chrono::steady_clock::now() + 200ms;
        if (cv.wait_until(lock, deadline, [] { return ready; }))
            std::cout << "condition met\n";
        else
            std::cout << "timed out, ready = " << std::boolalpha << ready << '\n';
    });
    std::this_thread::sleep_for(500ms);
    { std::lock_guard<std::mutex> lock(mtx); ready = true; }
    cv.notify_one();
    t.join();
}
timed out, ready = false

The notifier fires after 500 ms, well past the 200 ms deadline, so the waiter gives up. The later notify_one wakes no one and is harmless.

Two details matter here:

  • Use steady_clock for deadlines. system_clock can jump when the wall clock is adjusted, and a deadline based on it can then fire early or much too late.
  • Prefer wait_until when you loop. If your own code calls wait_for(lock, 1s, ...) repeatedly in an outer loop (for example to also check something else), every iteration restarts a fresh one-second timeout, so the total wait is unbounded. Computing one absolute deadline and calling wait_until keeps the total bounded. The predicate overloads of wait_for and wait_until handle spurious wakeups internally against the original deadline, so this only matters for loops you write yourself.

condition_variable_any and stop_token (C++20)

std::condition_variable_any works with any lock type that has lock() and unlock(), such as std::shared_lock or your own lock type. It is more general and typically has more overhead than std::condition_variable, so use it when you need the flexibility.

In C++20 it also gained wait overloads that take a std::stop_token. Combined with std::jthread, this gives a way to interrupt a waiting thread without adding your own stop flag to every predicate:

#include <chrono>
#include <condition_variable>
#include <iostream>
#include <mutex>
#include <queue>
#include <stop_token>
#include <thread>
using namespace std::chrono_literals;

std::mutex mtx;
std::condition_variable_any cv;
std::queue<int> jobs;

void worker(std::stop_token st) {
    std::unique_lock lock(mtx);
    while (true) {
        // Returns when the predicate is true OR a stop was requested.
        if (!cv.wait(lock, st, [] { return !jobs.empty(); })) {
            std::cout << "worker: stop requested, exiting\n";
            return;
        }
        int job = jobs.front();
        jobs.pop();
        std::cout << "worker: job " << job << '\n';
    }
}

int main() {
    std::jthread t(worker);
    { std::lock_guard lock(mtx); jobs.push(1); }
    cv.notify_one();
    std::this_thread::sleep_for(100ms);
    t.request_stop();  // wakes the waiter; ~jthread would also request stop and join
}
worker: job 1
worker: stop requested, exiting

The stop-aware wait returns the predicate’s value. false means it returned because of the stop request while the predicate was still false. This overload only exists on condition_variable_any, not on std::condition_variable.

When a condition variable is the wrong tool

  • One value produced once: std::promise/std::future already packages “wait for a result” and carries the value or exception with it.
  • Counting permits or a simple signal with no other state: std::counting_semaphore / std::binary_semaphore (C++20) do not need a separate mutex and predicate.
  • “Wait until N threads arrive”: std::latch (one-shot) or std::barrier (reusable) in C++20.
  • A flag that is only polled, never waited on: std::atomic<bool> is enough. C++20 also added atomic::wait/notify_one for blocking on a value change of a single atomic.

A condition variable is the general tool for “sleep until an arbitrary condition over shared data becomes true”. The specialized primitives are usually easier to get right when they fit.