Asio Deadlock Debugging: Async Callbacks, Locks, and Strands

The Asio Deadlock Problem

Boost.Asio async servers typically run with multiple threads calling io_context::run(). Completion handlers execute on whichever thread picks them up. This makes it tempting to protect shared state with mutexes — but it also creates a class of deadlocks that are timing-dependent and hard to reproduce.

The fundamental pattern:

  1. Thread A holds a mutex and waits for an async operation to complete
  2. The async operation’s completion handler, running on Thread B, tries to acquire the same mutex
  3. Thread B blocks forever — Thread A never releases the mutex because it’s waiting for Thread B

In synchronous code this would be obvious. In async code, the lock acquisition in step 1 and the handler in step 2 might be in completely different files.

Asio adds a second resource that people forget can run out: threads that are able to run handlers. A handler only runs when some thread inside io_context::run() picks it up (and, for a strand, when no other handler of that strand is running). Blocking one of those threads, or blocking inside a strand, removes capacity that the completion you are waiting for may need. Many “mutex deadlocks” in Asio code turn out on inspection to be this kind of starvation, with the mutex playing no real part. The patterns below separate the two cases, because the fixes differ.


Setting Up the Examples

All examples use Boost.Asio with C++17:

#include <boost/asio.hpp>
#include <mutex>
#include <condition_variable>
#include <thread>
#include <iostream>
#include <memory>

namespace asio = boost::asio;
using tcp = asio::ip::tcp;

Deadlock Pattern 1: Mutex Held While Waiting for Async Completion

This is the most common Asio deadlock:

class BrokenSession {
    tcp::socket socket_;
    std::mutex  mtx_;
    std::condition_variable cv_;
    bool        write_done_ = false;

public:
    // Called from inside another completion handler, i.e. on an io_context thread
    void sendSync(std::string data) {
        std::unique_lock<std::mutex> lock(mtx_);
        write_done_ = false;

        asio::async_write(socket_, asio::buffer(data),
            [this](boost::system::error_code ec, std::size_t) {
                // Must run on an io_context thread to wake the waiter below
                std::lock_guard<std::mutex> lk(mtx_);
                write_done_ = true;
                cv_.notify_one();
            });

        // cv_.wait releases mtx_ while waiting, so the mutex is NOT the problem.
        // The problem is the thread: this thread is an io_context thread, and
        // while it sleeps here it cannot run the completion handler.
        cv_.wait(lock, [this] { return write_done_; });   // hangs forever with 1 thread
    }
};

The mutex in this code is a red herring, which is exactly why the bug is confusing. condition_variable::wait releases the mutex while it sleeps, so the handler could acquire it without trouble. What it cannot get is a thread. With a single thread calling io_context::run(), and sendSync called from a handler on that thread, the only thread that could run the write’s completion handler is the one sleeping in cv_.wait. It waits for the handler, and the handler waits for it. With N io threads, the same code works until N handlers happen to block in sendSync at the same time, at which point every io thread is asleep and the server stops, which is why it tends to appear only under load.

The same happens with any blocking wait for async work on an io thread: std::future::get() on a future completed by a handler, use_future tokens waited on synchronously, or a semaphore released by a completion handler. Some code tries to escape by pumping the context manually (io_.poll_one() in a loop while waiting); that “works”, but it runs arbitrary other handlers in the middle of this one, re-entering code that was never written to be re-entrant, and it still deadlocks if those handlers block in the same way.

Fix: never block an io_context thread waiting for async completion. Instead, chain operations:

// CORRECT: chain — when write finishes, call the next step
class FixedSession : public std::enable_shared_from_this<FixedSession> {
    tcp::socket socket_;

public:
    void send(std::string data) {
        auto self = shared_from_this();
        auto buf = std::make_shared<std::string>(std::move(data));

        asio::async_write(socket_, asio::buffer(*buf),
            [self, buf](boost::system::error_code ec, std::size_t) {
                if (!ec) {
                    self->onWriteComplete();  // chain to next step
                }
            });
        // Return immediately — don't wait here
    }

    void onWriteComplete() {
        // Continue the session: read next request, process next message, etc.
        startRead();
    }
};

Chaining changes the shape of the code: “send, then continue” becomes “send, and continue in the completion handler”, and the function returns immediately. Two details keep it correct. The buffer is owned by a shared_ptr captured in the handler, because asio::buffer does not copy data and the caller’s string would otherwise be destroyed before the write finishes. And self keeps the session alive until the handler runs. If the logic genuinely reads better as sequential steps, C++20 coroutines (co_await asio::async_write(socket_, buf, asio::use_awaitable)) give you sequential-looking code that still never blocks a thread; the coroutine is suspended, and the thread goes back to running other handlers.

If a non-Asio thread (a UI thread, a legacy synchronous API) must wait for a result, that is acceptable, because that thread is not needed to run handlers. Waiting on a std::future there is fine; the rule is only about io_context threads.


Deadlock Pattern 2: Lock Order Inversion

Two threads take the same two mutexes but in opposite order:

std::mutex session_mutex;   // protects session state
std::mutex cache_mutex;     // protects a shared cache

// Thread A (handles incoming data):
void onReceive(const std::string& data) {
    std::lock_guard<std::mutex> session_lock(session_mutex);  // lock session FIRST
    // ... process data ...
    {
        std::lock_guard<std::mutex> cache_lock(cache_mutex);  // then lock cache
        cache.update(data);
    }
}

// Thread B (flushes cache periodically):
void flushCache() {
    std::lock_guard<std::mutex> cache_lock(cache_mutex);  // lock cache FIRST
    // ... flush cache data ...
    {
        std::lock_guard<std::mutex> session_lock(session_mutex);  // then lock session
        // Update session stats
    }
}

// DEADLOCK CYCLE:
// Thread A holds session_mutex, waits for cache_mutex
// Thread B holds cache_mutex, waits for session_mutex

This is not specific to Asio, but multi-threaded io_contexts make it much more likely, because receive handlers and timer handlers (like a periodic flushCache) run concurrently on different threads without any code explicitly starting threads. The window is small: both threads must take their first lock before either takes its second, so the deadlock may occur once in millions of calls, and never in a debugger.

Fix option 1: enforce a global lock order (always session → cache, never cache → session):

// Global rule: always acquire session_mutex before cache_mutex
void flushCache() {
    // Must acquire session_mutex first, even though we primarily want cache_mutex
    std::lock_guard<std::mutex> session_lock(session_mutex);
    std::lock_guard<std::mutex> cache_lock(cache_mutex);
    // ... flush ...
}

Fix option 2: acquire both atomically with std::scoped_lock (C++17):

// std::scoped_lock acquires multiple mutexes deadlock-free using a try-lock loop
void flushCache() {
    std::scoped_lock lock(session_mutex, cache_mutex);  // order doesn't matter
    // ... flush ...
}

void onReceive(const std::string& data) {
    std::scoped_lock lock(session_mutex, cache_mutex);
    // ... update both ...
}

std::scoped_lock with multiple arguments uses a deadlock-avoidance algorithm (similar to std::lock) that guarantees no cycle regardless of lock acquisition order.

The two fixes have different costs. A global order is free at runtime but only holds if everyone follows it; it breaks the day someone calls a function that takes session_mutex while holding cache_mutex, possibly through a callback several layers down. std::scoped_lock is robust against ordering mistakes, but only when both mutexes are known at the same point; it cannot help when the second lock is taken inside a function you call. It also means holding both locks for the whole scope. Often the better fix is structural: copy the data you need under one lock, release it, then take the other lock, so no code path ever holds two locks at once.


Deadlock Pattern 3: Strand Misuse

Strands serialize handlers for a connection. Deadlock occurs when you call a synchronous operation from inside a strand-serialized handler:

asio::strand<asio::io_context::executor_type> strand_;

void badHandler() {
    // This handler runs inside the strand
    // DON'T: block waiting for work that itself must run on the same strand
    std::promise<int> p;
    std::future<int> f = p.get_future();

    asio::post(strand_, [&p, this]() {
        p.set_value(doWork());   // queued on the strand behind badHandler()
    });

    int result = f.get();  // DEADLOCK: the strand cannot run the posted handler
                           // until badHandler() returns, and it never returns
}

A strand guarantees that at most one of its handlers runs at a time, on any thread. The posted lambda is queued behind the handler that is currently running, badHandler, so it cannot start until badHandler returns, and badHandler is waiting for it. Unlike the single-thread case in pattern 1, adding more io threads does not help here: the strand, not the thread pool, is the bottleneck. The same trap appears when a strand handler calls a “synchronous” helper that internally posts to the strand and waits, which is a common shape for adapters that wrap async APIs in blocking calls.

Fix: use asio::post instead of blocking gets, or structure the handler to return and let the chain continue:

// CORRECT: no blocking waits inside strand handlers
void goodHandler() {
    // Schedule the next step via post — returns immediately
    asio::post(strand_, [this]() {
        doWork();
    });
    // Return — let the strand execute doWork() after this handler finishes
}

The choice between post and dispatch matters here too. dispatch runs the function immediately if the caller is already executing on that strand (or executor), and otherwise queues it. That makes it faster, but it also means code after dispatch may run after the function has already completed, and if the caller holds a std::mutex that the function also locks, the immediate execution locks it a second time on the same thread, which is undefined behavior and in practice a self-deadlock. post always queues and never runs the function inside the caller, which is the safer default whenever locks or re-entrancy are involved.


Fixing with Per-Connection Strands

The strand-first approach eliminates most mutex needs for per-connection state:

class Session : public std::enable_shared_from_this<Session> {
    tcp::socket socket_;
    asio::strand<asio::any_io_executor> strand_;  // serializes this session's handlers

    // No mutex needed for these — strand guarantees serial access
    std::string write_buffer_;
    bool        writing_ = false;
    std::deque<std::string> pending_writes_;

public:
    Session(tcp::socket socket)
        : socket_(std::move(socket))
        , strand_(socket_.get_executor())
    {}

    // Can be called from any thread — always posts to strand
    void send(std::string data) {
        asio::post(strand_, [self = shared_from_this(), data = std::move(data)]() mutable {
            self->sendOnStrand(std::move(data));
        });
    }

private:
    // Always called on the strand — no mutex needed
    void sendOnStrand(std::string data) {
        pending_writes_.push_back(std::move(data));
        if (!writing_) {
            writeNext();
        }
    }

    void writeNext() {
        if (pending_writes_.empty()) {
            writing_ = false;
            return;
        }
        writing_ = true;
        write_buffer_ = std::move(pending_writes_.front());
        pending_writes_.pop_front();

        asio::async_write(socket_, asio::buffer(write_buffer_),
            asio::bind_executor(strand_,  // completion handler runs on strand too
                [self = shared_from_this()](boost::system::error_code ec, std::size_t) {
                    if (!ec) self->writeNext();
                }));
    }
};

This design handles concurrent sends safely with no locks:

  • send() can be called from any thread — it posts to the strand
  • sendOnStrand() and writeNext() run on the strand — no concurrent access
  • The write chain continues naturally without blocking

The queue is not only about thread safety. Asio forbids starting a second async_write on a socket while one is in progress, because the composed operation performs several partial writes and two of them interleaving would corrupt the stream. writing_ plus pending_writes_ enforce “one write in flight” and preserve message order. write_buffer_ is a member so the data stays alive until the write completes.

Every piece of this relies on all handlers of the session going through the strand. bind_executor makes the write completion run there; the read handlers need the same treatment, and a common oversight is a timer (steady_timer for idle timeouts) whose handler touches session state without being bound to the strand, which reintroduces a data race. A cleaner approach is to create the socket itself on a strand, for example by accepting with acceptor.async_accept(asio::make_strand(io), ...): then the socket’s default executor is the strand, and completion handlers use it without explicit bind_executor. (asio::executor, used in older examples, was replaced by any_io_executor in Boost 1.74.)


Debugging a Live Deadlock

Thread Dump with gdb

When a server hangs at ~0% CPU, attach gdb:

# Find the process ID
ps aux | grep my-server

# Attach and dump all thread backtraces
gdb -p <pid> -batch -ex "thread apply all bt full" -ex "quit" 2>&1 | tee deadlock.txt

# Or interactively:
gdb -p <pid>
(gdb) thread apply all bt full
(gdb) quit

Look for threads blocked in pthread_mutex_lock or std::condition_variable::wait. Find the cycle:

  • Thread 1: holds mutex A, waiting for mutex B
  • Thread 2: holds mutex B, waiting for mutex A (or waiting for a handler that needs mutex A)

The CPU reading is the first clue: a deadlock sits near 0% CPU, while a server at 100% on one core is more likely spinning in a loop or a livelock. In the backtraces, the useful question for each io thread is “why is this thread not inside io_context::run() waiting for work?” A healthy idle Asio thread is parked in epoll_wait (Linux) or GetQueuedCompletionStatus (Windows) under scheduler::run. A thread that is instead in __lll_lock_wait, pthread_cond_wait or std::future::get from inside a handler is part of the problem. With glibc, print *(pthread_mutex_t*)&mtx_ or looking at the mutex’s __owner field shows the thread ID (LWP) that holds a std::mutex, which lets you follow the chain from one thread to the next. Capture a core with gcore <pid> before restarting the process, so the analysis can continue after service is restored.

ThreadSanitizer for Lock-Order Inversion

TSan detects lock-order inversions before they cause a deadlock:

# Compile with ThreadSanitizer
clang++ -fsanitize=thread -g -O1 server.cpp -o server -lboost_system

# Run under stress — TSan logs order violations
./server

# TSan output example:
# WARNING: ThreadSanitizer: lock-order-inversion (potential deadlock)
# Cycle in lock order graph: M0 => M1 => M0

TSan’s value here is that it reports the potential deadlock: it records which locks were held when each lock was acquired, and flags a cycle in that graph even if the unlucky interleaving never happened during the run. One execution of each code path is enough, so a normal test suite finds lock-order inversions that stress tests would need hours to trigger. It does not detect pattern 1 or pattern 3, since those involve waiting for work rather than for a second mutex; for those, the backtrace and the “never block an io thread” rule are the tools. TSan slows the program considerably and increases memory use, so it belongs in a dedicated CI job rather than in production.

Watchdog Timer

Add a watchdog that logs if no progress is made in N seconds:

class Watchdog {
    asio::steady_timer timer_;
    std::chrono::seconds interval_;
    std::function<void()> callback_;
    std::atomic<int64_t> last_heartbeat_{0};

public:
    Watchdog(asio::io_context& io, std::chrono::seconds interval, std::function<void()> cb)
        : timer_(io), interval_(interval), callback_(std::move(cb))
    {
        arm();
    }

    void heartbeat() {
        last_heartbeat_.store(
            std::chrono::steady_clock::now().time_since_epoch().count());
    }

private:
    void arm() {
        timer_.expires_after(interval_);
        timer_.async_wait([this](boost::system::error_code ec) {
            if (!ec) {
                auto now = std::chrono::steady_clock::now().time_since_epoch().count();
                auto last = last_heartbeat_.load();
                if (now - last > interval_.count() * 1'000'000'000LL) {
                    callback_();   // log a warning or dump state
                }
                arm();
            }
        });
    }
};

// Usage:
Watchdog watchdog(io, std::chrono::seconds(30), []() {
    std::cerr << "WARNING: no progress for 30 seconds — possible deadlock\n";
    // dump active sessions, pending operations, etc.
});

One important caveat about where the watchdog runs: if it uses the same io_context as the server, and a deadlock blocks every io thread, the watchdog’s timer handler can never run either, so it stays silent in exactly the situation it exists for. Give it its own io_context on a dedicated thread (or a plain std::thread that sleeps and checks), and have the server’s handlers call heartbeat() as they make progress. Two smaller details in this sketch: the comparison assumes steady_clock ticks in nanoseconds, which is true for libstdc++ and libc++ but not guaranteed; storing std::chrono::duration_cast<std::chrono::nanoseconds>(...) makes it explicit. And the handler captures this, so the watchdog must outlive its timer or cancel it in its destructor.

When the watchdog fires, the most useful thing it can do is automatic: log the state it knows (sessions, queue lengths, last heartbeat time) and trigger a core dump or thread dump, because a hung production process is usually restarted before anyone can attach a debugger by hand.


Deadlock Prevention Checklist

RuleWhy
Never wait on an io thread or strand for a handler that needs that thread or strandThe handler can’t run — cycle
Use std::scoped_lock when acquiring multiple mutexesAtomic acquisition prevents order inversions
Prefer per-connection strands over mutexes for session stateStrand serialization without locks
Never block an io_context threadPrevents handlers from running — use post and chain instead
Document the lock hierarchyMakes order violations visible in code review
Run under TSan in CICatches order inversions before production

Why Asio programs deadlock, in short

  • The core pattern: blocking an io_context thread (or a strand handler) while waiting for a completion handler that needs that same thread or strand creates a cycle; the mutex is often incidental
  • Never block an io_context thread: it prevents completion handlers from running — use async chaining instead
  • Lock order inversion: two threads taking the same two mutexes in opposite order → use std::scoped_lock(m1, m2) for simultaneous acquisition
  • Strands: asio::strand serializes handlers without mutexes — prefer it for per-connection state
  • asio::bind_executor(strand_, handler): ensures the completion handler runs on the strand
  • Debug with gdb: thread apply all bt full reveals threads blocked in pthread_mutex_lock
  • TSan: -fsanitize=thread catches lock-order inversions at runtime before they deadlock in production
  • Watchdog timer: log a warning if no progress for N seconds — catches deadlocks in production

Frequently Asked Questions (FAQ)

Q. Does running io_context on a single thread make these deadlocks go away?

A. Only some of them. Lock-order inversion needs two threads holding locks, so it disappears with one io thread. Blocking while waiting for an async completion gets worse, though: if that single thread sits in cv.wait() or future.get(), the completion handler that would wake it can never run, so it hangs every time. The rule is the same either way: never block waiting for work that has to run on the thread you are blocking.