C++20 std::latch and std::barrier: Start Gates, Phased Work and Clean Exits

Key takeaways

std::latch is a single-use countdown for 'wait until N things have happened'; std::barrier is a reusable rendezvous for threads that work in lockstep phases. This article covers start gates, the barrier completion step, leaving a barrier with arrive_and_drop, and the count mistakes that turn both into hangs.

Before C++20, “wait until four workers have finished loading” meant a mutex, a counter, a std::condition_variable and a predicate loop, written by hand every time. C++20 added two small types in <latch> and <barrier> that cover the two most common shapes of that problem:

  • std::latch: a counter that goes from N down to zero exactly once. When it reaches zero, every current and future wait() returns immediately. It cannot be reset.
  • std::barrier: a rendezvous point for a fixed group of threads that work in phases. When all of them arrive, an optional completion function runs, the count resets, and everyone moves on to the next phase.

The rest of this article is about choosing between them and about the handful of mistakes that make them hang.

The examples compile with MSVC 2022 (/std:c++20). GCC needs version 11 or newer for these headers; GCC 10’s libstdc++ does not ship <latch> or <barrier> at all.

The latch interface in one screen

class latch {
public:
    static constexpr ptrdiff_t max() noexcept;
    constexpr explicit latch(ptrdiff_t expected);

    void count_down(ptrdiff_t update = 1);       // decrement, never blocks
    bool try_wait() const noexcept;              // true once the count is zero
    void wait() const;                           // block until the count is zero
    void arrive_and_wait(ptrdiff_t update = 1);  // count_down + wait
};

The key property is that the threads that count down and the threads that wait can be different. A latch isn’t tied to a group of participants. It just counts events. count_down() never blocks, so a worker can report “done” and carry on with something else.

Pattern 1: a start gate built from two latches

The most useful latch pattern I know uses two of them. One tells the main thread that every worker has finished its setup. The other releases all the workers at the same moment. It shows up in benchmarks (you don’t want thread creation time in the measurement) and in stress tests (you want the threads to actually collide).

#include <latch>
#include <thread>
#include <vector>
#include <atomic>
#include <cstdio>

int main() {
    constexpr int kWorkers = 4;
    std::latch ready(kWorkers);   // workers -> main: "I'm set up"
    std::latch go(1);             // main -> workers: "start now"
    std::atomic<long> total{0};

    std::vector<std::jthread> threads;
    for (int id = 0; id < kWorkers; ++id) {
        threads.emplace_back([&, id] {
            std::vector<int> local(1'000'000, id);   // per-thread setup
            ready.count_down();                      // report ready, don't block
            go.wait();                               // wait for the starting signal
            long sum = 0;
            for (int v : local) sum += v;
            total += sum;
        });
    }

    ready.wait();        // every worker finished its setup
    std::puts("all workers ready, starting");
    go.count_down();     // release all of them at once
    threads.clear();     // jthread joins in its destructor
    std::printf("total = %ld\n", total.load());
}

Output:

all workers ready, starting
total = 6000000

Two details matter. The workers call count_down() on ready, not arrive_and_wait(), because they need to block on go, not on each other. And go has a count of 1, so a single count_down() from main opens it for all of them. A latch with count 1 works as a one-shot “event” or “manual reset event” that never resets.

ready could also have been written with arrive_and_wait() if the workers only needed to wait for each other. That version makes the workers themselves the barrier for a single phase, which is exactly the case where a latch and a barrier overlap.

Counting down must not depend on the happy path

A latch doesn’t care why a count is missing. If one worker throws before its count_down(), ready.wait() blocks forever and the process hangs instead of crashing. That’s harder to diagnose. The fix is the same as for any resource release: put it in a destructor.

There is a second trap here. An exception that escapes a std::thread (or std::jthread) function calls std::terminate. Wrapping the thread construction in try/catch in main catches nothing thrown inside the thread. The worker has to catch its own exceptions and hand them back, for example as std::exception_ptr:

#include <latch>
#include <thread>
#include <vector>
#include <stdexcept>
#include <exception>
#include <cstdio>

class CountDownOnExit {
public:
    explicit CountDownOnExit(std::latch& l) : latch_(l) {}
    ~CountDownOnExit() { latch_.count_down(); }
    CountDownOnExit(const CountDownOnExit&) = delete;
    CountDownOnExit& operator=(const CountDownOnExit&) = delete;
private:
    std::latch& latch_;
};

void loadShard(int id) {
    if (id == 2) throw std::runtime_error("shard 2 is corrupt");
}

int main() {
    constexpr int kShards = 4;
    std::latch loaded(kShards);
    std::vector<std::exception_ptr> errors(kShards);

    std::vector<std::jthread> threads;
    for (int id = 0; id < kShards; ++id) {
        threads.emplace_back([&, id] {
            CountDownOnExit guard(loaded);   // runs on normal return and on unwind
            try {
                loadShard(id);
            } catch (...) {
                errors[id] = std::current_exception();  // an escaping exception would call std::terminate
            }
        });
    }

    loaded.wait();   // cannot hang: every thread counts down exactly once
    for (int id = 0; id < kShards; ++id) {
        if (!errors[id]) continue;
        try { std::rethrow_exception(errors[id]); }
        catch (const std::exception& e) { std::printf("shard %d failed: %s\n", id, e.what()); }
    }
}

Output:

shard 2 failed: shard 2 is corrupt

Each thread writes only its own slot in errors, and those writes happen before the count_down() in the guard’s destructor. count_down() synchronizes with the wait() that returns, so main can read the vector afterwards without a mutex.

The hang I have seen most often with latches wasn’t an exception, though. It was an early return. Someone added if (config.empty()) return; at the top of a worker, months after the latch count was written, and the service started hanging on startup only in environments where that config was missing. A guard object makes that edit harmless. A bare count_down() at the bottom of the function doesn’t.

Getting the count right

The count has to match the number of count_down() calls, not the number of threads. If one thread counts down twice (once per file it loads, say), the latch needs two counts for it. Being wrong in either direction is bad:

  • Too high: wait() never returns.
  • Too low: the latch opens early, and a later count_down() tries to go below zero. That violates the precondition update <= counter and is undefined behavior. The standard does not promise an exception, so don’t rely on one to catch the mistake.

The barrier interface

template<class CompletionFunction = /* unspecified no-op */>
class barrier {
public:
    using arrival_token = /* unspecified */;
    static constexpr ptrdiff_t max() noexcept;

    constexpr explicit barrier(ptrdiff_t expected, CompletionFunction f = CompletionFunction());

    [[nodiscard]] arrival_token arrive(ptrdiff_t update = 1);
    void wait(arrival_token&& token) const;
    void arrive_and_wait();
    void arrive_and_drop();
};

Where a latch counts events, a barrier coordinates participants. A fixed number of threads each arrive once per phase. When the last one arrives, the barrier:

  1. runs the completion function once, on one of the arriving threads,
  2. resets the count to the expected number of participants,
  3. releases everyone who is waiting.

Everything written before arrive in any thread is visible to the completion function, and everything the completion function writes is visible to every thread after its wait returns. That ordering is what makes the next example work without a single mutex or atomic.

Pattern 2: a phase loop with a completion step

A typical barrier workload is an iterative computation split across threads, where step k+1 needs every thread’s results from step k. Here is a small Jacobi iteration for a 1D heat problem. The rod is held at 100 on one end and 0 on the other, so the steady state is a straight line. Each thread updates its slice into the next buffer. The completion step swaps the buffers and decides whether to stop.

#include <barrier>
#include <thread>
#include <vector>
#include <cmath>
#include <cstdio>

int main() {
    constexpr int kThreads = 4;
    constexpr int kN = 66;
    std::vector<double> a(kN, 0.0), b(kN, 0.0);
    a.front() = b.front() = 100.0;          // fixed boundary values
    a.back()  = b.back()  = 0.0;

    std::vector<double>* cur = &a;
    std::vector<double>* next = &b;
    std::vector<double> maxDelta(kThreads, 0.0);  // one slot per thread, no locking
    int iterations = 0;
    bool done = false;

    // Runs exactly once per phase, after every thread has arrived and
    // before any of them is released.
    auto onPhaseComplete = [&]() noexcept {
        double d = 0.0;
        for (double x : maxDelta) d = std::fmax(d, x);
        std::swap(cur, next);
        ++iterations;
        done = (d < 1e-9) || iterations == 100000;
    };
    std::barrier sync(kThreads, onPhaseComplete);

    auto worker = [&](int t) {
        const int chunk = (kN - 2) / kThreads;
        const int lo = 1 + t * chunk;
        const int hi = (t == kThreads - 1) ? kN - 1 : lo + chunk;
        while (true) {
            const auto& in = *cur;
            auto& out = *next;
            double local = 0.0;
            for (int i = lo; i < hi; ++i) {
                out[i] = 0.5 * (in[i - 1] + in[i + 1]);
                local = std::fmax(local, std::fabs(out[i] - in[i]));
            }
            maxDelta[t] = local;
            sync.arrive_and_wait();   // completion step ran; cur/next/done are updated
            if (done) break;
        }
    };

    std::vector<std::jthread> threads;
    for (int t = 0; t < kThreads; ++t) threads.emplace_back(worker, t);
    threads.clear();

    std::printf("iterations: %d, value at middle: %.2f\n", iterations, (*cur)[kN / 2]);
}

Output:

iterations: 16106, value at middle: 49.23

The middle point is index 33 of 0..65, and the straight line gives 100 × (1 − 33/65) ≈ 49.23, so the result checks out. (Jacobi iteration converges slowly. The large iteration count is a property of the method, not of the barrier.)

Things this example gets right, and that are easy to get wrong:

  • Shared decisions belong in the completion function. done, cur and next are written in exactly one place, while every worker is parked. If each worker computed its own “should I stop?” from the partial data it saw, one thread could exit while the others go on to the next phase and wait for it forever.

  • No thread reads cur while another swaps it. The swap happens inside the completion step, and the workers only dereference the pointers after arrive_and_wait() returns.

  • The completion function is noexcept. The standard requires is_nothrow_invocable_v<CompletionFunction&>. Drop noexcept from the lambda and MSVC rejects it with:

    error C2338: static_assert failed: 'N4950 [thread.barrier.class]/5: is_nothrow_invocable_v<CompletionFunction&> shall be true'
  • The completion function is short. It runs while all the other threads are blocked, so any time spent in it is added to every phase of every thread. Reduce, swap, set a flag. Don’t do I/O there.

The mistake I’d warn about first with barriers is the “almost uniform” loop. Every thread runs for (i = 0; i < n; ++i) { ...; sync.arrive_and_wait(); } until someone gives one thread a continue that skips the barrier, or a break on a local condition. Now that thread has arrived one time fewer than the others, and the program deadlocks on the last phase, often only for certain inputs. When I review barrier code, I check that every path through the loop body hits the barrier exactly once and that the loop exit condition comes from shared state decided in the completion step, as above.

Leaving a barrier early: arrive_and_drop

Sometimes a participant really does have less work, for example the thread that owns a short tail of the data. It can’t just return, because the barrier would keep expecting it. arrive_and_drop() counts as its arrival for the current phase and lowers the expected count for every later phase:

#include <barrier>
#include <thread>
#include <vector>
#include <cstdio>

int main() {
    int phase = 0;
    std::barrier sync(3, [&]() noexcept { std::printf("-- phase %d done\n", phase++); });

    auto worker = [&](int id, int rounds) {
        for (int r = 0; r < rounds; ++r)
            sync.arrive_and_wait();
        std::printf("worker %d leaving after %d rounds\n", id, rounds);
        sync.arrive_and_drop();   // counts for this phase, shrinks later phases
    };

    std::vector<std::jthread> ts;
    ts.emplace_back(worker, 0, 1);
    ts.emplace_back(worker, 1, 3);
    ts.emplace_back(worker, 2, 3);
}

One run printed this (the order of the “leaving” lines within a phase can differ between runs):

-- phase 0 done
worker 0 leaving after 1 rounds
-- phase 1 done
-- phase 2 done
worker 2 leaving after 3 rounds
worker 1 leaving after 3 rounds
-- phase 3 done

Worker 0’s drop is its arrival for phase 1. From phase 2 on, the barrier waits for two threads. The final phase is completed by two drops, and the completion function still runs for it.

The same mechanism is how you make a barrier exception-safe. A guard whose destructor calls arrive_and_drop() stops a failing worker from stranding the rest. Whether the remaining threads can produce a meaningful result without it is a separate question. Usually they should also check an error flag set in the completion step.

Splitting arrive and wait

arrive_and_wait() is wait(arrive()). Splitting the two lets a thread announce “my part of this phase is done” and do independent work before blocking:

auto token = sync.arrive();     // my results for this phase are published
prefetchNextInput();            // doesn't touch data other threads are writing
sync.wait(std::move(token));    // now wait for the rest of the phase

The token belongs to the phase in which arrive() was called, and it must be passed to wait() on the same barrier. Anything done between the two calls must not touch data that the other threads or the completion function are still using for this phase.

Choosing between latch, barrier and condition_variable

You needUse
”Wait until N events have happened”, oncestd::latch
A one-shot signal from one thread to manystd::latch with count 1
The same group of threads meeting at the end of every stepstd::barrier
One piece of code that runs between steps while everyone is pausedbarrier completion function
A participant that finishes earlystd::barrier::arrive_and_drop()
A wait on an arbitrary condition, a count that must be reset, or a producer/consumer queuestd::mutex + std::condition_variable
A single result or exception from one taskstd::future / std::promise

A latch and a barrier are not faster versions of a condition variable in any general sense. Their implementations are free to use atomics and platform wait primitives, but the real benefit is that the counting, the predicate loop and the spurious-wakeup handling are no longer your code, so you can’t get them wrong. If the wait condition isn’t simply “a counter reached zero”, a condition variable is still the right tool.

Lifetime

Both objects must outlive every thread that uses them. In the examples above they are locals in main, and the threads are joined (by jthread’s destructor) before those locals go out of scope. The failure mode is a latch that is a local in a function that starts detached threads and returns after wait(). A thread that is still inside its count_down() call at that moment is touching a destroyed object. Join the threads, or give the latch a lifetime longer than theirs, for example via std::shared_ptr.