C++ std::future and std::promise: Passing Results Between Threads vs std::async
Key takeaways
std::promise is the write end and std::future the read end of a one-shot channel that carries a value or an exception between threads. The post covers std::async launch policies, the blocking destructor trap, timeouts that do not cancel work, shared_future, and patterns for parallel work.
What are future and promise?
std::future and std::promise are C++11 components for asynchronous programming. They provide a mechanism to retrieve results from tasks running in separate threads, enabling clean separation between producers and consumers.
Why are they needed?:
- Result Delivery: Get results from asynchronous tasks
- Exception Propagation: Safely propagate exceptions across threads
- Synchronization: Wait for task completion without busy-waiting
- Clean API: Better than raw threads with shared variables
// ❌ Without future/promise: Complex, error-prone
std::mutex mtx;
std::condition_variable cv;
int result;
bool ready = false;
void compute() {
int value = expensive_computation();
{
std::lock_guard<std::mutex> lock(mtx);
result = value;
ready = true;
}
cv.notify_one();
}
// ✅ With future/promise: Simple, safe
std::future<int> fut = std::async(expensive_computation);
int value = fut.get(); // Clean and simple
The manual version needs four shared variables and a specific protocol (lock, set, unlock, notify; on the other side, wait with a predicate) that every reader of the code must verify. It also has no way to report failure: if expensive_computation throws inside the thread, the exception escapes the thread function and std::terminate kills the process. A future bundles all of that into one object — a shared state holding either a value or an exception, plus the synchronization to wait for it — and get() either returns the value or rethrows the exception in the waiting thread.
(The examples below assume #include <iostream>, <thread>, <chrono>, and using namespace std; where they use unqualified names, to keep them short.)
std::async
Using std::async covered in Asynchronous Execution with async, you can execute tasks asynchronously and retrieve the result using future.
#include <future>
#include <iostream>
using namespace std;
int compute(int x) {
this_thread::sleep_for(chrono::seconds(1));
return x * x;
}
int main() {
// Asynchronous execution
future<int> result = async(launch::async, compute, 10);
cout << "Calculating..." << endl;
// Wait for the result
cout << "Result: " << result.get() << endl; // 100
}
How std::async works:
- Creates a task (function + arguments)
- Launches it asynchronously (new thread or deferred)
- Returns a
futureobject get()retrieves the result (blocks if not ready)
// Step-by-step execution
auto future = std::async(std::launch::async, compute, 10);
// 1. New thread created
// 2. compute(10) starts executing
// 3. Main thread continues
std::cout << "Doing other work...\n";
int result = future.get();
// 4. Blocks until compute() finishes
// 5. Returns result
Arguments passed to std::async (and std::thread) are copied into the task by default, like compute, 10 here. That protects against dangling references, but it means std::async(process, bigVector) copies the whole vector; pass std::cref(bigVector) to share it by reference — and then you are responsible for keeping it alive until the task finishes — or std::move(bigVector) to hand it over.
“New thread” is what the standard describes (launch::async runs the task “as if in a new thread”), but implementations differ: libstdc++ and libc++ create a real std::thread per call, while MSVC runs std::async tasks on the Windows thread pool. The observable difference is thread-local storage: on MSVC, thread_local variables may carry values over from a previous task that ran on the same pool thread.
promise and future
Relationship between promise and future
graph LR
A[promise] -->|get_future| B[future]
C[Producer Thread] -->|set_value| A
B -->|get| D[Consumer Thread]
style A fill:#e1f5ff
style B fill:#ffe1e1
style C fill:#e1ffe1
style D fill:#ffe1ff
void compute(promise<int> p, int x) {
this_thread::sleep_for(chrono::seconds(1));
p.set_value(x * x); // Set the result
}
int main() {
promise<int> p;
future<int> f = p.get_future();
thread t(compute, move(p), 10);
cout << "Calculating..." << endl;
cout << "Result: " << f.get() << endl; // 100
t.join();
}
The promise is moved into the thread because it is move-only — there is exactly one writer. get_future() must be called before the move (and only once; a second call throws future_error with future_already_retrieved). The promise/future pair is the low-level building block: std::async and std::packaged_task create one internally. You use a raw promise when the value is produced somewhere that is not a simple function return — a callback from a network library, an event that arrives on another thread, or a value computed halfway through a long-running worker that keeps going afterwards.
Two rules about the promise side cause most bugs. Setting a value twice throws future_error (promise_already_satisfied), which typically happens when both a success path and an error path fire. And if the promise is destroyed without ever being set — the worker returned early, or threw before set_value — the waiting get() throws future_error with broken_promise instead of hanging forever. That is a helpful failure, but only if someone catches it; the exception-handling version later in this article shows how to use set_exception so the consumer receives the real error rather than “broken promise”.
Workflow
sequenceDiagram
participant Main as Main Thread
participant Promise as promise
participant Future as future
participant Worker as Worker Thread
Main->>Promise: create
Main->>Future: get_future()
Main->>Worker: start thread
Main->>Main: other work
Worker->>Worker: compute
Main->>Future: get()
Note over Main,Future: waiting...
Worker->>Promise: set_value(result)
Promise->>Future: deliver result
Future->>Main: return result
Main->>Worker: join()
launch Policies
Comparison of Policies
| Policy | Execution Time | Thread | Overhead | Suitable Tasks |
|---|---|---|---|---|
| async | Immediate | New Thread | High | CPU-intensive, long tasks |
| deferred | On get() | Current Thread | Low | Short tasks, conditional execution |
| async|deferred | Implementation-defined | Automatic | Medium | General use |
// async: new thread
auto f1 = async(launch::async, compute, 10);
// deferred: delayed execution (on get() call)
auto f2 = async(launch::deferred, compute, 10);
// automatic selection
auto f3 = async(compute, 10);
The default policy (async | deferred) lets the implementation choose, and that freedom is a trap. If it picks deferred, the task does not run until someone calls get() or wait() — and never runs if nobody does. Code that polls f3.wait_for(0s) == future_status::ready in a loop then spins forever, because a deferred future reports future_status::deferred, not ready. Mainstream implementations currently pick async in practice, but the standard does not promise it. Scott Meyers’ advice in Effective Modern C++ is the common rule: pass launch::async explicitly whenever the task must actually run concurrently.
deferred is occasionally useful on purpose: it gives you a lazily evaluated value with the same future interface as real asynchronous work, so the caller does not need to know which one it received.
Parallel sums, downloads, timeouts, and exceptions
Splitting a sum across async tasks
#include <future>
#include <vector>
#include <numeric>
long long sumRange(int start, int end) {
long long sum = 0;
for (int i = start; i < end; i++) {
sum += i;
}
return sum;
}
int main() {
const int N = 1000000;
const int numThreads = 4;
const int chunkSize = N / numThreads;
vector<future<long long>> futures;
// Parallel execution
for (int i = 0; i < numThreads; i++) {
int start = i * chunkSize;
int end = (i == numThreads - 1) ? N : (i + 1) * chunkSize;
futures.push_back(async(launch::async, sumRange, start, end));
}
// Collect results
long long total = 0;
for (auto& f : futures) {
total += f.get();
}
cout << "Total: " << total << endl; // 499999500000
}
Two details in this example are easy to get wrong. The sum of 0 through 999,999 is 499,999,500,000, far beyond the range of a 32-bit int, so the accumulator and the future’s value type must be long long — with int, signed overflow is undefined behavior and the printed total is garbage. And when N is not divisible by the number of threads, N / numThreads silently drops the remainder, which is why the last chunk ends at N rather than at (i + 1) * chunkSize.
Collecting with f.get() in order is fine even though tasks finish in any order: each get() just waits for its own task. The number of tasks should be close to std::thread::hardware_concurrency() for CPU-bound work; launching one std::async per element creates thousands of threads on libstdc++, which is slower than a plain loop.
Simulated file downloads
#include <future>
#include <vector>
string downloadFile(const string& url) {
// Simulate download
this_thread::sleep_for(chrono::seconds(1));
return "Content from " + url;
}
int main() {
vector<string> urls = {
"http://example.com/file1",
"http://example.com/file2",
"http://example.com/file3"
};
vector<future<string>> futures;
// Parallel download
for (const auto& url : urls) {
futures.push_back(async(launch::async, downloadFile, url));
}
// Collect results
for (auto& f : futures) {
cout << f.get() << endl;
}
}
Waiting with a timeout
int longComputation() {
this_thread::sleep_for(chrono::seconds(5));
return 42;
}
int main() {
auto f = async(launch::async, longComputation);
// Wait for 2 seconds
if (f.wait_for(chrono::seconds(2)) == future_status::ready) {
cout << "Result: " << f.get() << endl;
} else {
cout << "Timeout" << endl;
}
}
wait_for only stops waiting; it does not stop the computation. After printing “Timeout”, main returns, the future f is destroyed — and because it came from std::async, its destructor blocks until longComputation finishes. The program still takes five seconds to exit. C++ futures have no cancellation mechanism; to abandon work, the task itself must check a flag, such as a std::atomic<bool> or C++20’s std::stop_token with std::jthread, and return early.
Exceptions crossing from worker to caller
int divide(int a, int b) {
if (b == 0) {
throw runtime_error("Division by zero");
}
return a / b;
}
int main() {
auto f = async(launch::async, divide, 10, 0);
try {
int result = f.get(); // Re-throws exception
cout << result << endl;
} catch (const exception& e) {
cout << "Error: " << e.what() << endl;
}
}
The exception thrown inside the worker thread is captured as a std::exception_ptr in the shared state and rethrown by get() in the calling thread, with its original type — catch (const runtime_error&) works. That is the main practical advantage over a bare std::thread, where an exception escaping the thread function calls std::terminate.
shared_future
int compute() {
this_thread::sleep_for(chrono::seconds(1));
return 42;
}
int main() {
shared_future<int> sf = async(launch::async, compute).share();
// Accessible from multiple threads
thread t1([sf]() {
cout << "Thread 1: " << sf.get() << endl;
});
thread t2([sf]() {
cout << "Thread 2: " << sf.get() << endl;
});
t1.join();
t2.join();
}
share() converts the future into a shared_future and leaves the original empty (valid() returns false). A shared_future is copyable, and each thread should hold its own copy, as the lambdas above do by capturing sf by value; calling get() concurrently on copies is safe, but calling it concurrently on the same shared_future object from several threads is a data race. shared_future::get() returns a const T& instead of moving the value out, which is why it can be called repeatedly — and why the value must be safe to read from several threads at once. A typical use is a one-time initialization result (configuration, a loaded model) that many workers wait for.
get() twice, blocking destructors, and lost exceptions
Calling get() more than once
// ❌ get() can only be called once
future<int> f = async(compute, 10);
int x = f.get();
// int y = f.get(); // Throws exception
// ✅ Store the result
int result = f.get();
The blocking destructor of an async future
// ❌ Future destruction causes blocking
{
auto f = async(launch::async, compute, 10);
} // Blocks here
// ❌ Worse: the temporary future is destroyed at the semicolon
async(launch::async, task1); // blocks until task1 finishes
async(launch::async, task2); // then runs task2 — nothing is concurrent
// ✅ Keep the futures alive, and wait where you intend to
auto f1 = async(launch::async, task1);
auto f2 = async(launch::async, task2);
f1.get();
f2.get();
Only futures created by std::async have a blocking destructor; futures from a promise or packaged_task do not. The rule exists so that an async task can never outlive the variables it references, but it produces the second form above — “fire and forget” calls that silently run one after another. Compilers help a little: std::async is marked [[nodiscard]] in current standard libraries, so discarding the future produces a warning. If you really need fire-and-forget, use a std::thread that you detach() (and accept the lifetime risks) or a proper thread pool.
Exceptions nobody retrieves
// ❌ Ignoring exceptions
auto f = async(launch::async, []() {
throw runtime_error("Error");
});
// If f.get() is not called, the exception is ignored
// ✅ Handle exceptions
try {
f.get();
} catch (const exception& e) {
cout << e.what() << endl;
}
Advanced promise Usage
void compute(promise<int> p, int x) {
try {
if (x < 0) {
throw invalid_argument("Negative value not allowed");
}
p.set_value(x * x);
} catch (...) {
p.set_exception(current_exception());
}
}
int main() {
promise<int> p;
future<int> f = p.get_future();
thread t(compute, move(p), -10);
try {
cout << f.get() << endl;
} catch (const exception& e) {
cout << "Error: " << e.what() << endl;
}
t.join();
}
Running tasks concurrently, aggregating, and falling back
Several independent tasks in flight
class TaskPipeline {
std::vector<std::future<void>> tasks_;
public:
template<typename F>
void addTask(F&& task) {
tasks_.push_back(std::async(std::launch::async, std::forward<F>(task)));
}
void waitAll() {
for (auto& task : tasks_) {
task.wait();
}
}
~TaskPipeline() {
waitAll(); // Ensure all tasks complete
}
};
// Usage
TaskPipeline pipeline;
pipeline.addTask([]() { processData(); });
pipeline.addTask([]() { sendNotification(); });
pipeline.addTask([]() { updateDatabase(); });
pipeline.waitAll();
Despite the name, these tasks run concurrently, not as a pipeline — updateDatabase may finish before processData starts. wait() also discards exceptions: if sendNotification throws, the exception stays in its future and is silently dropped when the future is destroyed. Calling get() on each future in waitAll surfaces it (collect the first exception and rethrow after waiting for the rest, so no task is left running). The destructor’s waitAll() is redundant with the blocking destructors of std::async futures, but it documents the intent.
Aggregating results from many futures
template<typename T, typename F>
std::vector<T> parallelMap(const std::vector<T>& input, F func) {
std::vector<std::future<T>> futures;
for (const auto& item : input) {
futures.push_back(std::async(std::launch::async, func, item));
}
std::vector<T> results;
for (auto& future : futures) {
results.push_back(future.get());
}
return results;
}
// Usage
std::vector<int> numbers = {1, 2, 3, 4, 5};
auto squares = parallelMap(numbers, [](int x) { return x * x; });
Taking the callable as a template parameter F rather than std::function<T(const T&)> is not just an optimization: with the std::function signature, this call does not compile, because template argument deduction cannot deduce T from a lambda (a lambda is not a std::function) and fails with “no matching function for call to ‘parallelMap’” / “mismatched types”. It also avoids type erasure on every call. As in the parallel-sum example, one thread per element is only reasonable for a handful of expensive items; for large inputs, split into chunks or use C++17’s std::transform(std::execution::par, ...).
A timeout with a fallback value
template<typename T>
T getWithTimeout(std::future<T>& future,
std::chrono::milliseconds timeout,
T fallback) {
if (future.wait_for(timeout) == std::future_status::ready) {
return future.get();
}
return fallback;
}
// Usage
auto future = std::async(std::launch::async, expensiveComputation);
int result = getWithTimeout(future, std::chrono::seconds(5), -1);
if (result == -1) {
std::cout << "Timeout, using fallback\n";
}
The same caveat as the timeout example above applies: the fallback is returned, but expensiveComputation keeps running, and destroying future later blocks until it ends. Using -1 as both the fallback and a signal is also fragile if -1 is a legitimate result; std::optional<T> makes “no result in time” unambiguous.
FAQ
Q1: async vs thread?
A:
- async: Simple API, automatic result delivery, exception propagation
- thread: Fine-grained control, manual synchronization, no direct result return
// async: Simple
auto future = std::async(compute, 10);
int result = future.get();
// thread: Manual
int result;
std::thread t([&result]() { result = compute(10); });
t.join();
Choose async for tasks that return results. Choose thread for long-running background tasks.
Q2: When should I use future?
A:
- Asynchronous tasks: I/O operations, network requests
- Parallel computation: CPU-intensive tasks that can run concurrently
- Result delivery: When you need to return values from threads
- Exception handling: When you need safe exception propagation
// Perfect for async I/O
auto file1 = std::async(readFile, "data1.txt");
auto file2 = std::async(readFile, "data2.txt");
auto data1 = file1.get();
auto data2 = file2.get();
Q3: What about performance?
A: Creating a thread typically costs on the order of tens of microseconds on Linux (more on some platforms), plus the synchronization in the shared state. For tasks that take only microseconds, that overhead outweighs the benefit, so batch small work items into larger chunks.
// ❌ Bad: Overhead > computation
for (int i = 0; i < 1000; ++i) {
auto f = std::async([i]() { return i * 2; }); // Too much overhead
}
// ✅ Good: Batch processing
const int CHUNK_SIZE = 250;
std::vector<std::future<int>> futures;
for (int i = 0; i < 4; ++i) {
futures.push_back(std::async([i]() {
int sum = 0;
for (int j = i * CHUNK_SIZE; j < (i + 1) * CHUNK_SIZE; ++j) {
sum += j * 2;
}
return sum;
}));
}
Q4: Can future be reused?
A: No. get() moves the result out and can only be called once. Use shared_future for multiple accesses.
// ❌ future: Single access
std::future<int> f = std::async(compute);
int x = f.get();
// int y = f.get(); // Throws std::future_error
// ✅ shared_future: Multiple accesses
std::shared_future<int> sf = std::async(compute).share();
int x = sf.get();
int y = sf.get(); // OK
Q5: How do I handle timeouts?
A: Use wait_for() or wait_until() to check if the result is ready.
auto future = std::async(longComputation);
// wait_for: Relative timeout
if (future.wait_for(std::chrono::seconds(5)) == std::future_status::ready) {
int result = future.get();
} else {
std::cout << "Timeout\n";
}
// wait_until: Absolute timeout
auto deadline = std::chrono::system_clock::now() + std::chrono::seconds(5);
if (future.wait_until(deadline) == std::future_status::ready) {
int result = future.get();
}
Q6: What happens if I don’t call get()?
A: The destructor blocks until the task completes. This can cause unexpected blocking.
// ❌ Destructor blocks
{
auto f = std::async(std::launch::async, longTask);
} // Blocks here until longTask completes!
// ✅ Explicit handling
auto f = std::async(std::launch::async, longTask);
f.wait(); // Explicit wait
Q7: How do exceptions work with future?
A: Exceptions are stored in the future and re-thrown when get() is called.
auto future = std::async([]() {
throw std::runtime_error("Error!");
return 42;
});
try {
int result = future.get(); // Re-throws exception
} catch (const std::exception& e) {
std::cout << "Caught: " << e.what() << '\n';
}
Q8: What’s the difference between launch::async and launch::deferred?
A:
- launch::async: Starts right away, as if on a new thread (a thread pool on MSVC)
- launch::deferred: Runs lazily when
get()orwait()is called (in the calling thread)
// async: Immediate execution
auto f1 = std::async(std::launch::async, compute);
// compute() starts running NOW in new thread
// deferred: Lazy execution
auto f2 = std::async(std::launch::deferred, compute);
// compute() doesn't run yet
int result = f2.get(); // NOW compute() runs in current thread
Q9: Any resources for learning future/promise?
A:
- “C++ Concurrency in Action” (2nd Edition) by Anthony Williams
- cppreference.com - std::future
- cppreference.com - std::promise
- “Effective Modern C++” by Scott Meyers
Related Posts: Asynchronous Execution with async, shared_future, Thread Basics, packaged_task.
One-line Summary: std::future and std::promise provide a clean mechanism for asynchronous result delivery and exception propagation across threads.
Related Articles
- std::async and Launch Policies
- C++ std::thread basics — join/detach mistakes and mutex
- C++ packaged_task