C++20 Coroutines: promise_type, co_await Awaiters, Generators, and Task Types

Key takeaways

C++20 coroutines: coroutine_handle, promise_type, co_await, generators, Task patterns, and compiler support for async-style control flow.

C++20 coroutines are one of the most powerful — and most misunderstood — features the language has shipped in a decade. Unlike coroutines in Python, JavaScript, or Kotlin, C++ does not give you a ready-made Task or async/await runtime. Instead, the standard defines a set of compiler hooks (co_await, co_yield, co_return) and a customization protocol (promise_type) that you implement to get a usable abstraction. This is deliberate — it lets coroutines back generators, fibers, async I/O, and cooperative schedulers with zero forced runtime cost — but it also means the first time most C++ developers touch coroutines, they are actually writing a tiny compiler-facing state machine framework, not just calling an async API. That gap between “the feature looks small” and “the implementation surface is large” is the single biggest reason coroutines get a reputation for being hard.

Coroutine basics: what co_return actually triggers

#include <coroutine>
#include <iostream>
using namespace std;
struct Task {
    struct promise_type {
        Task get_return_object() {
            return Task{handle_type::from_promise(*this)};
        }
        
        suspend_never initial_suspend() { return {}; }
        suspend_never final_suspend() noexcept { return {}; }
        void return_void() {}
        void unhandled_exception() {}
    };
    
    using handle_type = coroutine_handle<promise_type>;
    handle_type coro;
    
    Task(handle_type h) : coro(h) {}
    ~Task() { if (coro) coro.destroy(); }
};
Task simpleCoroutine() {
    cout << "Coroutine start" << endl;
    co_return;
}
int main() {
    simpleCoroutine();
}

The moment the compiler sees any co_await, co_yield, or co_return in a function body, that function stops being an ordinary function. The compiler rewrites it into a state machine: local variables that need to survive a suspension point get hoisted into a heap-allocated coroutine frame, the function body is split at every suspension point into resumable chunks, and a promise_type object is created inside that frame to mediate every transition. Task::promise_type above is the minimum viable implementation — get_return_object() builds the object the caller actually sees, initial_suspend()/final_suspend() decide whether the coroutine starts and ends eagerly or lazily, and unhandled_exception() is where an exception thrown inside the coroutine body lands if you don’t rethrow it yourself.

The gotcha that trips up nearly everyone on their first pass: initial_suspend() returning suspend_never means the coroutine body starts running synchronously, on the calling thread, the instant you call simpleCoroutine() — the “Coroutine start” line prints before simpleCoroutine() even returns. If you instead return suspend_always, the body doesn’t run at all until something calls .resume() on the handle. Getting this backwards is a common source of “why did my supposedly async function block the caller” bugs, and it’s not a runtime bug at all — it’s a design decision baked into your promise_type that most tutorials gloss over.

Generators: cooperative iteration through co_yield

#include <coroutine>
#include <iostream>
using namespace std;
template<typename T>
struct Generator {
    struct promise_type {
        T current_value;
        
        Generator get_return_object() {
            return Generator{handle_type::from_promise(*this)};
        }
        
        suspend_always initial_suspend() { return {}; }
        suspend_always final_suspend() noexcept { return {}; }
        
        suspend_always yield_value(T value) {
            current_value = value;
            return {};
        }
        
        void return_void() {}
        void unhandled_exception() {}
    };
    
    using handle_type = coroutine_handle<promise_type>;
    handle_type coro;
    
    Generator(handle_type h) : coro(h) {}
    ~Generator() { if (coro) coro.destroy(); }
    
    bool move_next() {
        coro.resume();
        return !coro.done();
    }
    
    T current_value() {
        return coro.promise().current_value;
    }
};
Generator<int> counter(int max) {
    for (int i = 0; i < max; i++) {
        co_yield i;
    }
}
int main() {
    auto gen = counter(5);
    
    while (gen.move_next()) {
        cout << gen.current_value() << " ";  // 0 1 2 3 4
    }
}

co_yield i desugars into co_await promise.yield_value(i), which is why yield_value returns an awaiter (suspend_always here) rather than just storing the value and returning void. Every call to yield_value produces a suspension point: control returns to whoever called .resume(), the loop variable i and everything else live in the coroutine frame is frozen in place, and the frame stays allocated until either the generator is resumed again or coro.destroy() runs. This is exactly why generators are cheap compared to building an equivalent iterator by hand with an explicit state enum — the compiler generates that state machine for you from what reads like a plain for loop.

The part worth calling out because it’s genuinely easy to miss: Generator has no comparison-based end() sentinel and no operator*/operator++, so it isn’t a drop-in range in this minimal form — you drive it explicitly with move_next()/current_value(). Wiring it up to work with range-based for loops requires an Iterator wrapper that calls move_next() from operator++ and coro.done() from operator==. That’s boilerplate every hand-rolled generator type reimplements slightly differently, which is precisely the gap std::generator (added in C++23, not C++20) and libraries like cppcoro’s generator<T> were built to close — if your toolchain supports std::generator, prefer it over writing your own promise_type for simple cases; reach for a manual implementation only when you need a custom allocator, custom exception propagation, or a shape std::generator doesn’t support.

co_await and the awaiter protocol

struct Awaitable {
    bool await_ready() { return false; }
    void await_suspend(coroutine_handle<>) {}
    void await_resume() {}
};
Task asyncFunction() {
    cout << "Start" << endl;
    co_await Awaitable{};
    cout << "Resumed" << endl;
}

co_await expr compiles down to a fixed three-step protocol against an awaiter object (obtained from expr directly, or via an operator co_await, or via the promise’s await_transform): first await_ready() is checked — if it returns true, the coroutine proceeds immediately without ever suspending; if false, the coroutine suspends and await_suspend(handle) is called with a handle to the now-suspended coroutine, so the awaiter can store that handle and resume it later from anywhere, including another thread; finally, once resumed, await_resume() runs and its return value becomes the value of the whole co_await expression. The trivial Awaitable above always suspends and does nothing interesting in await_suspend, so it’s really only useful for demonstrating the mechanism — a real awaiter registers handle with an event loop, a timer, or an I/O completion callback and calls .resume() on it once that external event fires.

This is also where the sharpest edge in the whole feature lives: await_suspend runs after the coroutine has already suspended, which means the coroutine frame — including any local variables — is in a frozen, incomplete state at that exact moment. If await_suspend throws, or if it resumes the handle on another thread while the calling thread is still executing code that assumed the coroutine hadn’t suspended yet, you get exactly the kind of order-of-operations bug that doesn’t reproduce reliably and doesn’t show up as an obvious crash near the code you’d suspect.

Fibonacci, Range, File-Line, and Timer Coroutines

Example 1: Fibonacci generator

Generator<int> fibonacci(int n) {
    int a = 0, b = 1;
    
    for (int i = 0; i < n; i++) {
        co_yield a;
        int temp = a;
        a = b;
        b = temp + b;
    }
}
int main() {
    auto fib = fibonacci(10);
    
    while (fib.move_next()) {
        cout << fib.current_value() << " ";
    }
}

Example 2: Range generator

Generator<int> range(int start, int end, int step = 1) {
    for (int i = start; i < end; i += step) {
        co_yield i;
    }
}
int main() {
    for (auto gen = range(0, 10, 2); gen.move_next();) {
        cout << gen.current_value() << " ";  // 0 2 4 6 8
    }
}

Both examples show the actual selling point of coroutine-based generators: a, b, temp, and the loop counter i all persist across every co_yield, without you writing a single explicit state field. Compare that to the pre-C++20 idiom of hand-writing an iterator class that manually stores a and b as member variables and re-derives the “current step” logic inside operator++ — the coroutine version is shorter and, more importantly, harder to get subtly wrong, because the control flow in the source matches the control flow at runtime.

Example 3: Read file lines

#include <fstream>
Generator<string> readLines(const string& filename) {
    ifstream file(filename);
    string line;
    
    while (getline(file, line)) {
        co_yield line;
    }
}
int main() {
    auto lines = readLines("input.txt");
    
    while (lines.move_next()) {
        cout << lines.current_value() << endl;
    }
}

This one deserves a second look precisely because it looks so innocent. file and line are locals inside the coroutine, so they live in the coroutine frame and survive suspension correctly — that part is fine. The trap is what happens if you refactor readLines to take const string& filename and the caller passes a temporary (readLines("input.txt"s + suffix)), or if you change the signature to capture a reference to some object that lives on the caller’s stack. The coroutine frame is heap-allocated and can genuinely outlive the calling function’s stack frame, but a const string& parameter is not automatically copied into that heap frame — only the reference is. If the referred-to object is destroyed before the generator is fully consumed, every subsequent .move_next() reads through a dangling reference. This is arguably the single most common category of coroutine bug in production code: parameters passed by reference into a coroutine are not safe by default the way they are in an ordinary function, because “ordinary function” implies the call stack outlives the call, and a coroutine breaks that assumption the instant it suspends.

Example 4: Async timer

#include <chrono>
#include <thread>
struct Timer {
    chrono::milliseconds duration;
    
    bool await_ready() { return false; }
    
    void await_suspend(coroutine_handle<> h) {
        thread([h, d = duration]() {
            this_thread::sleep_for(d);
            h.resume();
        }).detach();
    }
    
    void await_resume() {}
};
Task asyncTask() {
    cout << "Start" << endl;
    
    co_await Timer{chrono::seconds(1)};
    cout << "After 1 second" << endl;
    
    co_await Timer{chrono::seconds(1)};
    cout << "After 2 seconds total" << endl;
}

Timer is a minimal but honest illustration of cross-thread resumption: await_suspend spawns a detached thread that sleeps and then calls h.resume() from that thread, not the original one. This is the pattern every real async I/O library follows in spirit — register a continuation, return control to the caller, and resume later from whatever context the underlying event actually completes on (an I/O completion port thread, a timer thread, a network reactor’s poll loop). The .detach() call is doing real, structural work here: because the thread outlives the function that spawned it, nothing about the Timer object’s own lifetime blocks program shutdown, but it also means you have zero built-in cancellation — if asyncTask’s Task gets destroyed while the sleeping thread is still in flight, that thread will still call .resume() on a handle whose coroutine frame has already been destroyed. Production async frameworks solve this with explicit cancellation tokens or by keeping the coroutine frame alive (via reference counting) until every outstanding continuation has either fired or been cancelled; this toy Timer does neither, which is fine for a demo and actively dangerous if copied verbatim into real code.

sequenceDiagram
    participant Caller
    participant Coroutine as asyncTask (frame)
    participant Awaiter as Timer::await_suspend
    participant Thread as detached thread

    Caller->>Coroutine: call asyncTask()
    Coroutine->>Coroutine: print "Start"
    Coroutine->>Awaiter: co_await Timer{1s}
    Awaiter->>Thread: spawn + detach, capture handle h
    Awaiter-->>Caller: control returns (coroutine suspended)
    Thread->>Thread: sleep_for(1s)
    Thread->>Coroutine: h.resume()
    Coroutine->>Coroutine: print "After 1 second"
    Coroutine->>Awaiter: co_await Timer{1s}
    Note over Coroutine,Thread: same pattern repeats,\nno cancellation path exists

Coroutines vs threads

// Thread (heavier)
void threadExample() {
    thread t([] {
        // work
    });
    t.join();
}
// Coroutine (lighter)
Task coroutineExample() {
    co_await someOperation();
    // work
}

Differences:

  • Thread: an OS-scheduled resource with its own stack (typically megabytes reserved, though usually not all committed) and a real context-switch cost enforced by the kernel scheduler.
  • Coroutine: a user-space state machine with a heap-allocated frame sized to exactly what it needs, resumed and suspended entirely by library or application code with no kernel involvement.

The practical decision is rarely “coroutines are strictly better” — it’s about what you’re modeling. If the work is CPU-bound and you want it to run in parallel on multiple cores, you need threads (or a thread pool underneath your coroutines); coroutines by themselves don’t grant parallelism, only concurrency. Where they win decisively is I/O-bound code with many outstanding operations at once: a server juggling ten thousand open connections cannot afford ten thousand OS threads, but it can afford ten thousand suspended coroutine frames, because a suspended coroutine costs only the memory its frame occupies, not a kernel-scheduled stack and a context-switch slot.

Why coroutines are hard to debug

This is worth its own section because it’s the complaint I hear most often from teams adopting coroutines, and it’s not a skill-issue complaint — the tooling genuinely lags behind the language feature. A stack trace captured inside a coroutine after it has been resumed does not show the original call chain that first invoked it; it shows whatever resumed it this time, which might be an event loop’s dispatch function, a thread pool worker, or another coroutine’s await_suspend. The logical “who called this” and the physical “what’s on the call stack right now” diverge completely once a coroutine has suspended and resumed at least once. Add symmetric transfer (returning another coroutine’s handle from await_suspend to avoid stack growth across chained co_awaits) and the picture gets worse — the compiler can tail-call directly into the next coroutine’s resume point, so even the “resumer” in the stack trace can be misleading.

I’ve watched engineers spend the better part of a day chasing what looked like heap corruption in a coroutine-heavy service, only to find the actual bug was a Task object destructing (and calling coro.destroy()) while a Timer-style awaiter still had an outstanding .resume() scheduled from a background thread — almost exactly the cancellation gap the Example 4 timer has above. The crash reported a garbled stack inside std::coroutine_handle<>::resume, which is technically accurate and completely unhelpful: the actual defect was a lifetime ordering issue between the coroutine frame’s destruction and a dangling continuation, three logical steps removed from where the crash surfaced. GDB and LLDB have added some coroutine-frame-aware unwinding since, but the fundamental issue remains — anywhere a coroutine can be resumed from a different logical call site than the one that suspended it, “read the stack trace” stops being a reliable debugging strategy, and you end up adding explicit logging around every initial_suspend, final_suspend, and await_suspend just to reconstruct the actual sequence of events.

Missing promise_type, Leaked Handles, and Dangling References

Problem 1: Missing promise_type

// Error without promise_type
Task myCoroutine() {
    co_return;
}
// Define promise_type inside Task
struct Task {
    struct promise_type {
        // ...
    };
};

The compiler looks up promise_type via std::coroutine_traits<Task>::promise_type, which by default resolves to a nested Task::promise_type. If it’s missing, or your return type is a plain type without one, you get a compile error pointing at the coroutine’s return type rather than at the missing nested type — a confusing first encounter for anyone who hasn’t seen this mechanism before.

Problem 2: Forgetting to destroy the handle

// Leak if destructor omits destroy
Generator<int> gen = counter(10);
// Destructor should call coro.destroy()
~Generator() {
    if (coro) coro.destroy();
}

Because the coroutine frame is heap-allocated (via operator new, which you can override per-promise-type to plug in a custom allocator or arena), nothing frees it automatically when the wrapper object holding the handle goes out of scope — unlike a std::unique_ptr, coroutine_handle is a non-owning, copyable raw handle by design. If your RAII wrapper’s destructor doesn’t call .destroy() on every path (including exception-unwinding paths), the frame — and anything it holds, including captured objects and open resources like the ifstream in Example 3 — leaks silently. This is easy to miss because leaks in coroutine frames don’t show up as an obviously wrong result; the program just slowly grows its heap footprint until something else notices.

Problem 3: Local variable lifetime and dangling references across suspension points

Locals declared inside the coroutine body itself are safely hoisted into the coroutine frame — the compiler handles that correctly and automatically. The danger is anything the coroutine doesn’t own outright: reference parameters, pointers, iterators, or views (std::string_view, std::span) into caller-owned storage. None of those are copied into the frame just because they’re used inside a coroutine; only the coroutine’s own locals are. If the referred-to object’s lifetime doesn’t outlive every suspension point the coroutine can reach, you have a use-after-free that often only manifests intermittently, depending on scheduling — exactly the kind of bug that passes code review and passes most unit tests, then shows up under production load when timing shifts expose the gap. The safe default is to capture by value into the coroutine’s own parameters (pass std::string instead of const std::string& when the coroutine’s lifetime isn’t guaranteed to be nested inside the caller’s) and to be explicit about which values genuinely need to survive across a co_await or co_yield.

I ran into a version of exactly this on a project that wrapped a legacy synchronous API in a coroutine-based Task for a request pipeline: a lambda captured a reference to a request-scoped object by reference ([&request]) inside a helper that then got wrapped in a coroutine, on the (false) assumption that the coroutine would run to completion before the request object went out of scope, the same assumption that holds for an ordinary function call. It didn’t — the coroutine suspended waiting on a downstream call, the request object was torn down at the end of what looked like its natural scope, and the coroutine resumed later against freed memory. AddressSanitizer caught it immediately once we had a repro, but the failure in production had shown up first as an intermittent, unrelated-looking corruption in a completely different field of the response struct, because the freed memory had already been reused for something else by the time the dangling read happened. The fix was mechanical — capture request by value, or wrap it in a shared_ptr kept alive by the coroutine itself — but finding it took longer than it should have precisely because the crash site and the actual defect were separated by a suspension point.

Coroutine state and the handle interface

coroutine_handle<> h = ...;
h.resume();
h.done();
h.destroy();
h.promise();

coroutine_handle<> (type-erased, no access to the promise) and coroutine_handle<Promise> (typed, exposes .promise()) are intentionally thin, non-owning views into a frame that already exists — creating one doesn’t allocate anything, and copying one just copies a pointer. .resume() continues execution from the last suspension point until the next suspension point, final_suspend, or a co_return/fall-off-the-end. .done() is only well-defined to check after the coroutine has reached its final suspend point; calling it on a coroutine that hasn’t started or is mid-execution is undefined behavior, not a false/true answer — a distinction that’s easy to miss when skimming the reference documentation. .destroy() runs the frame’s destructors and frees its storage; calling .resume() on a handle after .destroy() (or double-destroying a handle) is, again, undefined behavior with no runtime check, because these primitives are deliberately as close to zero-overhead as the language can make them — all the safety has to come from the wrapper types you build on top, which is exactly why Task and Generator exist in the first place rather than exposing raw handles to callers.

FAQ

Q1: When should I use coroutines?

A: Async I/O, generators, state machines, cooperative multitasking.

Q2: Coroutines vs async/await in other languages?

A: C++ coroutines are low-level; libraries like cppcoro add ergonomics.

Q3: Performance?

A: Much lighter than OS threads; many concurrent coroutines are feasible, though each still costs a heap allocation for its frame unless HALO elides it.

Q4: Compiler support?

A: GCC 10+, Clang 14+, MSVC 2019+.

Q5: Debugging?

A: Harder than plain functions; stack traces after a resume don’t reflect the original call chain, so add explicit logging around suspend/resume points and use a coroutine-aware debugger build where available.