std::async and Launch Policies: When async Actually Runs on Another Thread

Key takeaways

std::launch::async guarantees the task runs as if on a new thread; deferred runs it lazily inside get() or wait(). The default lets the library choose, the returned future blocks in its destructor, and there is no way to cancel a running task.

std::async is the shortest way in standard C++ to run a function and get its result back later:

#include <future>

int square(int x) { return x * x; }

int main() {
    std::future<int> f = std::async(square, 10);
    int r = f.get();   // 100
}

That simplicity hides several decisions: whether a thread is created at all, when the function runs, what happens if you never call get(), and how exceptions get back to you. Getting any of these wrong produces code that works in tests and quietly loses its concurrency, or hangs, in real use. The examples below were compiled with g++ 10 (-std=c++20) and the behavior shown is what they did.

The three launch policies

std::async(policy, f, args...) takes one of:

  • std::launch::async — run f as if on a new thread, starting now.
  • std::launch::deferred — do not run anything yet. f runs on the thread that first calls get() or wait() on the future, synchronously, inside that call.
  • std::launch::async | std::launch::deferred — the implementation picks. This is what you get when you omit the policy.

A quick way to see the difference is to compare thread IDs:

#include <future>
#include <iostream>
#include <thread>

int main() {
    auto main_id = std::this_thread::get_id();
    auto fa = std::async(std::launch::async,    [&] { return std::this_thread::get_id() == main_id; });
    auto fd = std::async(std::launch::deferred, [&] { return std::this_thread::get_id() == main_id; });
    std::cout << std::boolalpha
              << "async on main thread: "    << fa.get() << '\n'    // false
              << "deferred on main thread: " << fd.get() << '\n';   // true
}

A deferred task that nobody waits on never runs at all. That makes deferred useful for “compute this only if someone needs it” and nearly useless for concurrency.

What “as if on a new thread” means in practice

With libstdc++ (GCC), launch::async creates a fresh std::thread per call; there is no pool. Starting a few long-running tasks this way is fine, but calling std::async for thousands of small work items creates thousands of threads, and if the system refuses to create one, std::async throws std::system_error (with resource_unavailable_try_again).

MSVC’s implementation runs launch::async tasks on the Windows thread pool instead. That is cheaper, but it is observable: thread_local variables may keep their values from a previous task that ran on the same pooled thread, which a real “new thread” would not do. If your task relies on fresh thread-local state, that difference matters (see thread_local pitfalls).

The default policy and the wait_for trap

Because the default policy allows deferred execution, code that polls a future can hang:

auto f = std::async(work);                       // default policy
while (f.wait_for(std::chrono::milliseconds(10)) != std::future_status::ready) {
    do_other_things();                           // may loop forever
}

If the library chose deferred, wait_for returns std::future_status::deferred immediately every time, and the task never starts because nobody called get() or wait(). In the GCC build used here, the default policy ran tasks asynchronously, so the loop happened to work — which is exactly how this bug survives testing and appears on another platform, or under load when an implementation decides not to spawn another thread.

This is the std::async failure I consider most dangerous, because the code reads correctly. Two fixes: pass std::launch::async explicitly whenever you actually want concurrency, or handle the deferred status:

if (f.wait_for(std::chrono::seconds(0)) == std::future_status::deferred) {
    f.wait();   // run it now, on this thread
}

My rule is to never omit the policy in code that depends on the task running in parallel. The default is only appropriate when you genuinely do not care whether the work happens now, later, or inline.

The future’s destructor blocks

A std::future obtained from std::async is special: if it is the last reference to a shared state that is not ready yet, its destructor waits for the task to finish. Futures from std::promise or std::packaged_task do not do this.

The classic consequence is fire-and-forget code that turns out to be sequential:

std::async(std::launch::async, [] { std::this_thread::sleep_for(300ms); });
std::async(std::launch::async, [] { std::this_thread::sleep_for(300ms); });
// Took at least 600 ms: each temporary future blocked at the end of its statement

C++20 marks std::async as [[nodiscard]], and GCC 10 warns:

warning: ignoring return value of '...std::async(std::launch, _Fn&&, _Args&& ...)...', declared with attribute 'nodiscard' [-Wunused-result]

Storing the future in a variable moves the wait to the end of that variable’s scope; it does not remove it. The same rule makes timeouts less useful than they look:

{
    auto slow = std::async(std::launch::async, [] {
        std::this_thread::sleep_for(500ms);
        return 42;
    });
    if (slow.wait_for(100ms) == std::future_status::timeout) {
        std::cout << "timed out, leaving scope\n";
    }
}   // blocks here until the 500 ms task finishes

wait_for tells you the task is not done, but you cannot abandon it: there is no cancellation, and leaving the scope waits anyway. For a real timeout, the task itself must be interruptible, for example by checking a std::atomic<bool> or a C++20 std::stop_token (see std::jthread and stop_token), and even then you should treat “timed out” as “asked it to stop”, not “it stopped”.

Arguments are copied

Like std::thread, std::async decay-copies its arguments into the task. Passing a large container by value copies it; passing something to a function that takes a non-const reference does not compile unless you wrap it:

void bump(int& x) { ++x; }

int x = 0;
std::async(std::launch::async, bump, std::ref(x)).get();   // x == 1
// std::async(std::launch::async, bump, x);                // does not compile

For read-only large data, pass std::cref(data), and make sure the data outlives the task. A lambda capturing by reference has the same lifetime requirement; the destructor-blocking behavior protects you only as long as the future does not outlive the referenced objects.

Exceptions travel through the future

If the function throws, the exception is stored in the shared state and rethrown from get():

int divide(int a, int b) {
    if (b == 0) throw std::runtime_error("division by zero");
    return a / b;
}

auto f = std::async(std::launch::async, divide, 10, 0);
try {
    f.get();
} catch (const std::exception& e) {
    std::cout << "caught: " << e.what() << '\n';   // caught: division by zero
}

get() consumes the state: afterwards f.valid() is false. Calling get() again is undefined behavior according to the standard; libstdc++ happens to throw std::future_error with “No associated state”. If several consumers need the result, convert with f.share() to a std::shared_future, whose get() can be called repeatedly and from multiple threads (each thread should use its own copy of the shared_future).

If you never call get(), a stored exception is silently discarded when the future is destroyed. Code that only calls wait() on each future should still call get() somewhere to surface failures.

Waiting while holding a lock

Because get() and wait() may run a deferred task inline, or block until an async one finishes, calling them while holding a mutex is a classic way to deadlock:

std::mutex m;
auto f = std::async([&] {            // default policy
    std::lock_guard lock(m);         // the task needs m
    return compute();
});
std::lock_guard lock(m);
int r = f.get();                     // deferred: runs the task here, which locks m again

With launch::deferred, get() runs the lambda on the current thread, which already owns m; locking a non-recursive std::mutex twice from one thread is undefined behavior and in practice usually hangs. With launch::async, the task runs on another thread, blocks on m, and get() waits for it forever: a plain two-party deadlock. The default policy picks one of the two, so the program hangs either way, just for different reasons. The rule is the same as for joining threads: release locks before waiting on a future, or make sure the task never needs them.

The same inline execution affects thread_local variables. A deferred task sees the caller’s thread-local values, because it runs on the caller’s thread, while an async one sees a new or pooled thread’s values. Code that stores per-request state in thread_local behaves differently depending on which policy the library chose.

Splitting work across a few tasks

For a small, fixed number of CPU-bound chunks, std::async with an explicit policy is a reasonable tool:

int sum_range(const std::vector<int>& d, std::size_t b, std::size_t e) {
    return std::accumulate(d.begin() + b, d.begin() + e, 0);
}

int parallel_sum(const std::vector<int>& data, std::size_t n) {
    std::size_t chunk = data.size() / n;
    std::vector<std::future<int>> parts;
    for (std::size_t i = 0; i < n; ++i) {
        std::size_t b = i * chunk;
        std::size_t e = (i + 1 == n) ? data.size() : b + chunk;   // last chunk takes the remainder
        parts.push_back(std::async(std::launch::async, sum_range, std::cref(data), b, e));
    }
    int total = 0;
    for (auto& f : parts) total += f.get();
    return total;
}
// 1,000,001 ones split into 4 chunks -> 1000001

Collecting results in order with get() is fine here because every chunk has to finish anyway. Whether this is faster than a single loop depends on the data size and the machine; for plain sums the thread start-up cost can dominate small inputs, and std::reduce with std::execution::par expresses the same idea without manual chunking.

When to use something else

  • Many small tasks, or tasks that should not each cost a thread: a thread pool, with futures from std::packaged_task.
  • A result produced by code that is not a single function call, such as a callback or an event loop: std::promise.
  • Long-lived background workers that need to be stopped: std::jthread with a stop_token.
  • Waiting for “something happened” rather than “a value is ready”: std::condition_variable.

std::async is at its best for a handful of independent computations whose results you will definitely collect in the same scope. Outside that shape, its implicit waits and missing cancellation tend to work against you.