Production WebSocket in C++: Handshake Validation, Frames, Heartbeats and Backpressure with Beast

Introduction: “My WebSocket keeps dropping”

Scenario 1: disconnect every ~30 seconds

// ❌ NAT / firewall evicts idle TCP flows
// If the client is silent for longer than the middlebox timer,
// routers may RST the socket.
ws_.async_read(buffer_, [](beast::error_code ec, std::size_t) {
    // ec == connection_reset or connection_aborted
});

Why? NATs and firewalls track sessions and reap idle TCP entries. WebSocket can be quiet for a long time while still logically “open,” so the path looks dead. Mitigation: send Ping/Pong (or app-level keepalives) every 20–30s—shorter than the smallest idle timeout on the path.

The frustrating part of this failure is that neither end is told when the middlebox forgets the connection. The mapping simply disappears; the next packet either gets a RST from a device that no longer recognizes the flow or is silently dropped, in which case the sender only notices after TCP retransmissions give up, which can take many minutes. Load balancers are the usual culprit in practice: AWS ALB and many reverse proxies close connections idle for about 60 seconds by default, and nginx’s proxy_read_timeout defaults to 60 s as well. The “every 30 seconds” or “every 60 seconds” pattern in the logs is the fingerprint of such a timer. TCP keepalive does not help here with default settings, because the Linux default waits two hours before the first probe.

Scenario 2: HTTP 400 on handshake

// ❌ Server returns 400 Bad Request
ws_.async_handshake(host, "/chat",
    [](beast::error_code ec) {
        // ec == bad_request
        // server log: "Missing Sec-WebSocket-Key"
    });

Causes: missing Sec-WebSocket-Key, bad Upgrade, wrong version—anything that violates RFC 6455. On the client, Beast reports any non-101 response as websocket::error::upgrade_declined; the actual status code is only visible if you pass a response_type to async_handshake and inspect it. When Beast is on both ends, a 400 usually means something in between (a proxy that strips Upgrade/Connection hop-by-hop headers, or an HTTP/1.0 hop) rewrote the request.

Scenario 3: huge messages exhaust RAM

// ❌ Accepting a 100 MiB frame blows the buffer
ws_.async_read(buffer_, [](beast::error_code ec, std::size_t bytes) {
    // bytes == 100 * 1024 * 1024
});

Mitigation: always set read_message_max. Beast’s default is 16 MiB per message, which is generous; tune per workload.

A single oversized message is not the real threat; the threat is many connections each sending a message just under the limit. With a 16 MiB cap and 1,000 connected clients, a hostile or buggy client population can make the server buffer gigabytes. Set the cap to the largest legitimate message your protocol allows, not to “big enough to never hit”.

More scenarios

Scenario 4: random WSS drops under load (multithreaded server)

If multiple threads touch the same websocket::stream without a strand, reads and writes interleave inside the TLS layer, producing corrupted records or frames. Browsers then close the connection with a protocol or TLS error, which looks like “random drops” that only appear under load and only with several io_context threads—serialize every operation on a stream through its strand.

Scenario 5: reconnect storms

Immediate reconnect loops hammer the server. Use exponential backoff and jitter.

Scenario 6: broadcast write storms

Calling async_write for thousands of sessions at once floods the executor—use per-session queues and backpressure.

Goals:

  • Byte-level handshake understanding
  • Frames with concrete Text/Binary/Ping/Pong/Close examples
  • Full heartbeat design
  • Error catalog with fixes
  • Best practices: reconnect, backpressure, strands
  • Production: metrics, graceful shutdown Prerequisites: Boost.Beast 1.70+, C++17.

Mental model

Treat sockets as addresses and async I/O as scheduled delivery—strands keep a single connection’s handlers ordered.


Ops-focused: these patterns come from production C++ services, not toy echo servers.

Handshake anatomy

HTTP upgrade request (client → server)

Every WebSocket begins as an HTTP Upgrade request. Example aligned with RFC 6455:

GET /chat HTTP/1.1
Host: example.com:8080
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==
Sec-WebSocket-Version: 13
Origin: https://example.com

Header cheat sheet:

HeaderRequiredNotes
Upgrade: websocket✅Request protocol switch
Connection: Upgrade✅HTTP upgrade hop
Sec-WebSocket-Key✅16 random bytes → Base64 (mitigates proxy cache tricks)
Sec-WebSocket-Version: 13✅Only standardized version
OriginRecommendedBrowser CORS checks
Sec-WebSocket-ProtocolOptionalNegotiate subprotocols (chat, json, …)

Generating Sec-WebSocket-Key

#include <random>
#include <boost/beast/core/detail/base64.hpp>
// RFC 6455: 16 random bytes → Base64
std::string generate_websocket_key() {
    std::random_device rd;
    std::mt19937 gen(rd());
    std::uniform_int_distribution<> dis(0, 255);
    
    unsigned char key[16];
    for (int i = 0; i < 16; ++i) {
        key[i] = static_cast<unsigned char>(dis(gen));
    }
    
    std::string result;
    result.resize(boost::beast::detail::base64::encoded_size(16));
    result.resize(boost::beast::detail::base64::encode(
        &result[0], key, 16));
    
    return result;
}

Server response (101 Switching Protocols)

HTTP/1.1 101 Switching Protocols
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Accept: s3pPLMBiTxaQ9kYGzzhZRbK+xOo=

Computing Sec-WebSocket-Accept

#include <openssl/sha.h>
#include <boost/beast/core/detail/base64.hpp>
std::string compute_accept(const std::string& key) {
    const std::string magic = "258EAFA5-E914-47DA-95CA-C5AB0DC85B11";
    std::string input = key + magic;
    
    unsigned char hash[SHA_DIGEST_LENGTH];
    SHA1(reinterpret_cast<const unsigned char*>(input.data()),
         input.size(), hash);
    
    std::string result;
    result.resize(boost::beast::detail::base64::encoded_size(SHA_DIGEST_LENGTH));
    result.resize(boost::beast::detail::base64::encode(
        &result[0], hash, SHA_DIGEST_LENGTH));
    
    return result;
}

Algorithm: SHA1(Sec-WebSocket-Key + "258EAFA5-E914-47DA-95CA-C5AB0DC85B11") → Base64

The key/accept exchange is not authentication and not security; the GUID is public and anyone can compute the answer. Its purpose is to prove that the server actually understood the WebSocket handshake, rather than being an HTTP server or cache that echoed headers back, and to stop a browser from being tricked into speaking WebSocket to a server that never agreed to it. Likewise the key only needs to be unpredictable enough to be a nonce, which is why std::mt19937 is acceptable here. Both helpers use boost::beast::detail::base64, which is an internal namespace that can change between Boost releases; that is fine for a demonstration, but in real code you never need these functions, since async_handshake and async_accept generate and verify the headers for you. Write them by hand only when implementing or debugging a WebSocket stack.

Handshake sequence

sequenceDiagram
    participant C as Client
    participant S as Server
    
    C->>S: TCP connect
    C->>S: HTTP GET + Upgrade + Sec-WebSocket-Key
    S->>S: Validate key, compute Accept
    S->>C: HTTP 101 + Sec-WebSocket-Accept
    Note over C,S: WebSocket established
    C->>S: WebSocket frames
    S->>C: WebSocket frames

Handshake failure modes

StatusTypical cause
400 Bad RequestMissing Sec-WebSocket-Key, bad Upgrade
403 ForbiddenOrigin check failed
426 Upgrade RequiredWrong Sec-WebSocket-Version
503 Service Unavailableoverload / connection cap

Frames with worked examples

Frame layout (RFC 6455)

graph LR
    subgraph Header
        A[FIN 1bit] --> B[RSV 3bit]
        B --> C[Opcode 4bit]
        C --> D[Mask 1bit]
        D --> E[Payload Len 7bit]
    end
    E --> F[Extended 0/2/8 byte]
    F --> G[Mask Key 0/4 byte]
    G --> H[Payload Data]

Opcode reference

OpcodeValueMeaningDirection
Continuation0x0Continues previous fragmentBoth
Text0x1UTF-8 textBoth
Binary0x2Binary payloadBoth
Close0x8Close connectionBoth
Ping0x9Heartbeat probeBoth
Pong0xAHeartbeat replyBoth

Text frame (masked client → server)

Client → server: “Hello” (5 bytes)

Byte 0: 0x81 (FIN=1, Opcode=0x1 Text)
Byte 1: 0x85 (Mask=1, Payload Len=5)
Bytes 2-5: Masking Key (4 random bytes)
Bytes 6-10: "Hello" XOR Masking Key
// Masking required for client → server
void mask_payload(uint8_t* data, size_t len, const uint8_t key[4]) {
    for (size_t i = 0; i < len; ++i) {
        data[i] ^= key[i % 4];
    }
}

Why mask: mitigates cache poisoning when broken intermediaries mis-classify traffic.

The masking key is chosen fresh for every frame, so a malicious script cannot control the bytes that actually appear on the wire. Without it, a page could craft a WebSocket payload that looks like an HTTP request to a transparent proxy that does not understand WebSocket, and poison that proxy’s cache. Masking is therefore required only from client to server: a server must close the connection (code 1002) if it receives an unmasked client frame, and a client must fail if it receives a masked one. It is not encryption; the key travels in the frame header. Beast applies and removes masks automatically according to the stream’s role.

The payload length encoding also explains a common hand-parsing bug: 0–125 fits in the 7-bit field, 126 means “the next 2 bytes hold the length” and 127 means “the next 8 bytes do”, in network byte order. Parsers that forget the byte-order conversion work for small test messages and break on the first message over 125 bytes.

Ping frame

Byte 0: 0x89 (FIN=1, Opcode=0x9 Ping)
Byte 1: 0x00 (MASK=0 server→client, len=0)

If payload present, Pong echoes it.

Pong frame

Byte 0: 0x8A (FIN=1, Opcode=0xA Pong)
Byte 1: 0x00 (Payload Len=0)

Close frame

Byte 0: 0x88 (FIN=1, Opcode=0x8 Close)
Byte 1: 0x02 (Payload Len=2)
Bytes 2-3: close code (e.g. 1000 normal, 1001 going away, 1002 protocol error)
Bytes 4+: optional UTF-8 reason

Common close codes:

CodeMeaning
1000Normal Closure
1001Going away (server shutdown, etc.)
1002Protocol Error
1003Unsupported Data
1006Abnormal closure (no close frame)
1007Invalid payload (UTF-8)
1011Internal Error

1006 is special: it is never sent on the wire. It is what an API reports locally when the connection ended without a close frame, such as after a network drop or a process crash. Seeing many 1006 closes in browser logs means connections are dying rather than being closed, which points back at idle timeouts and heartbeats. Control frames (Close, Ping, Pong) may carry at most 125 bytes of payload and cannot be fragmented, which is why a close reason is limited to 123 bytes after the 2-byte code.

Handling control frames in Beast

ws_.control_callback(
    [](websocket::frame_type kind, beast::string_view) {
        switch (kind) {
            case websocket::frame_type::ping:
                // Beast sends Pong automatically
                break;
            case websocket::frame_type::pong:
                // observe heartbeat reply
                break;
            case websocket::frame_type::close:
                // peer initiated close
                break;
        }
    });

The control callback is invoked from inside a pending read operation: Beast processes control frames while it is reading, answers pings, and handles the close handshake, then calls your callback for observation. Two consequences follow. If no async_read is outstanding, incoming pings are not answered and pongs are not observed, so a session that only writes (a pure push feed) must still keep a read pending. And the callback must not start another operation on the stream synchronously; if you need to react, post the work to the stream’s executor.

Ping/Pong heartbeat

Ping/Pong sequence

sequenceDiagram
    participant C as Client
    participant S as Server
    
    loop Every 30s
        C->>S: Ping
        S->>C: Pong (auto)
    end
    
    Note over C: No Pong within 10s
    C->>C: Treat as dead → reconnect

Client: send Ping + Pong timeout

class WebSocketClientWithHeartbeat
    : public std::enable_shared_from_this<WebSocketClientWithHeartbeat> {
    websocket::stream<beast::tcp_stream> ws_;
    beast::flat_buffer buffer_;
    net::steady_timer ping_timer_;
    net::steady_timer pong_timeout_;
    bool pong_received_ = false;
    
public:
    explicit WebSocketClientWithHeartbeat(net::io_context& ioc)
        : ws_(net::make_strand(ioc)),
          ping_timer_(ws_.get_executor()),
          pong_timeout_(ws_.get_executor()) {}
    
    void start_heartbeat() {
        pong_received_ = true;
        schedule_ping();
    }
    
private:
    void schedule_ping() {
        ping_timer_.expires_after(std::chrono::seconds(30));
        ping_timer_.async_wait(
            [self = shared_from_this()](beast::error_code ec) {
                if (ec) return;
                self->send_ping();
            });
    }
    
    void send_ping() {
        pong_received_ = false;
        pong_timeout_.expires_after(std::chrono::seconds(10));
        pong_timeout_.async_wait(
            [self = shared_from_this()](beast::error_code ec) {
                if (ec) return;
                if (!self->pong_received_) {
                    std::cerr << "Pong timeout - reconnecting\n";
                    self->reconnect();
                    return;
                }
            });
        
        ws_.async_ping({},
            [self = shared_from_this()](beast::error_code ec) {
                if (ec) {
                    std::cerr << "Ping failed: " << ec.message() << "\n";
                    return;
                }
                self->schedule_ping();
            });
    }
    
    void on_pong() {
        pong_received_ = true;
        pong_timeout_.cancel();
    }
    
    void reconnect() {
        // reconnect with exponential backoff
    }
};

This class shows the moving parts, but note what it leaves unwired: nothing calls on_pong(). It has to be hooked up through control_callback (as in the server example below), and a read must be pending for that callback to run. The two timers run on the stream’s strand (they are constructed from ws_.get_executor()), so pong_received_ is only touched from one logical thread and needs no atomic. Beast allows a ping to be sent while a message write is in progress, because control frames are interleaved between the frames of a data message, but only one ping, pong or close may be outstanding at a time.

Before writing this by hand, check whether Beast’s built-in timeout option already does what you need. websocket::stream_base::timeout has an idle_timeout and a keep_alive_pings flag. With keep-alive pings enabled, Beast sends a ping when half the idle timeout has passed without incoming data and closes the connection if the full timeout passes with nothing received, so a dead peer is detected without any timers of your own. timeout::suggested(role_type::server) turns this on with a 300-second idle timeout, which is too long to beat a 60-second load balancer; setting idle_timeout to about 40 seconds gives a ping roughly every 20 seconds. A hand-written heartbeat is still useful when you need application-level liveness, for example checking that the peer’s application responds, not just its WebSocket layer.

Server: auto Pong to Ping

Beast answers Ping with Pong by default. Custom logging example:

ws_.control_callback(
    [self = shared_from_this()](
        websocket::frame_type kind, beast::string_view payload) {
        if (kind == websocket::frame_type::ping) {
            // Beast sends Pong automatically
            // manual: ws_.async_pong(payload);
        } else if (kind == websocket::frame_type::pong) {
            // client Pong after server-initiated Ping
            self->on_pong_received();
        }
    });

Server → client Ping (optional)

Servers may initiate Ping to verify the peer is still alive.

void server_send_ping() {
    ws_.async_ping("heartbeat",
        [self = shared_from_this()](beast::error_code ec) {
            if (ec) {
                // write failure ⇒ dead connection
                self->close_session();
            }
        });
}

A successful async_ping completion only means the frame was handed to the operating system, not that the peer received it. A write to a dead connection often succeeds for a while because the data sits in the kernel’s send buffer. That is why the client above pairs every ping with a pong timeout; a failed ping is a sure sign of a dead connection, but a successful one proves nothing.

Beast examples

Async client (handshake + read + ping)

#include <boost/beast.hpp>
#include <boost/asio.hpp>
#include <iostream>
namespace beast = boost::beast;
namespace websocket = beast::websocket;
namespace net = boost::asio;
using tcp = net::ip::tcp;
class CompleteWebSocketClient
    : public std::enable_shared_from_this<CompleteWebSocketClient> {
    websocket::stream<beast::tcp_stream> ws_;
    beast::flat_buffer buffer_;
    net::steady_timer ping_timer_;
    std::string host_;
    std::string path_;
    
public:
    explicit CompleteWebSocketClient(net::io_context& ioc)
        : ws_(net::make_strand(ioc)),
          ping_timer_(ws_.get_executor()) {}
    
    void connect(const std::string& host, const std::string& port,
                 const std::string& path = "/") {
        host_ = host;
        path_ = path;
        
        tcp::resolver resolver(ws_.get_executor());
        resolver.async_resolve(host, port,
            beast::bind_front_handler(&CompleteWebSocketClient::on_resolve,
                                     shared_from_this()));
    }
    
private:
    void on_resolve(beast::error_code ec,
                    tcp::resolver::results_type results) {
        if (ec) {
            std::cerr << "Resolve: " << ec.message() << "\n";
            return;
        }
        
        beast::get_lowest_layer(ws_).async_connect(results,
            beast::bind_front_handler(&CompleteWebSocketClient::on_connect,
                                     shared_from_this()));
    }
    
    void on_connect(beast::error_code ec,
                    tcp::resolver::results_type::endpoint_type ep) {
        if (ec) {
            std::cerr << "Connect: " << ec.message() << "\n";
            return;
        }
        
        ws_.async_handshake(host_, path_,
            beast::bind_front_handler(&CompleteWebSocketClient::on_handshake,
                                     shared_from_this()));
    }
    
    void on_handshake(beast::error_code ec) {
        if (ec) {
            std::cerr << "Handshake: " << ec.message() << "\n";
            return;
        }
        
        std::cout << "WebSocket connected\n";
        do_read();
        start_ping();
    }
    
    void do_read() {
        ws_.async_read(buffer_,
            beast::bind_front_handler(&CompleteWebSocketClient::on_read,
                                     shared_from_this()));
    }
    
    void on_read(beast::error_code ec, std::size_t bytes) {
        if (ec) {
            if (ec != websocket::error::closed) {
                std::cerr << "Read: " << ec.message() << "\n";
            }
            return;
        }
        
        std::cout << "Received: "
                  << beast::buffers_to_string(buffer_.data()) << "\n";
        buffer_.consume(buffer_.size());
        do_read();
    }
    
    void start_ping() {
        ping_timer_.expires_after(std::chrono::seconds(30));
        ping_timer_.async_wait(
            [self = shared_from_this()](beast::error_code ec) {
                if (ec) return;
                self->ws_.async_ping({},
                    [self](beast::error_code ec) {
                        if (!ec) self->start_ping();
                    });
            });
    }
};

The client is a chain of completion handlers, each holding a shared_ptr to the object through bind_front_handler(..., shared_from_this()). That is what keeps the object alive while an operation is pending: when the last handler finishes without starting another operation (for example after an error), the reference count drops to zero and the client is destroyed. This is also why connect() must be called on an object already owned by a shared_ptr; calling shared_from_this() on a stack object throws std::bad_weak_ptr. The stream is constructed on net::make_strand(ioc), so every handler for this connection runs serialized even when several threads call ioc.run().

Two production gaps remain. The tcp::resolver is a local variable, but it must outlive async_resolve; in Asio, destroying the resolver cancels the pending operation, so it belongs in a member. And the handshake is sent immediately after connect without setting a connect or handshake timeout; beast::get_lowest_layer(ws_).expires_after(...) before async_connect, and the websocket timeout option after the handshake, keep a hung server from stalling the client forever.

Async echo server

class CompleteWebSocketSession
    : public std::enable_shared_from_this<CompleteWebSocketSession> {
    websocket::stream<beast::tcp_stream> ws_;
    beast::flat_buffer buffer_;
    
public:
    explicit CompleteWebSocketSession(tcp::socket socket)
        : ws_(std::move(socket)) {}
    
    void run() {
        ws_.set_option(websocket::stream_base::timeout::suggested(
            beast::role_type::server));
        ws_.read_message_max(64 * 1024);  // 64 KiB cap
        
        ws_.async_accept(
            beast::bind_front_handler(&CompleteWebSocketSession::on_accept,
                                     shared_from_this()));
    }
    
private:
    void on_accept(beast::error_code ec) {
        if (ec) {
            std::cerr << "Accept: " << ec.message() << "\n";
            return;
        }
        
        do_read();
    }
    
    void do_read() {
        ws_.async_read(buffer_,
            beast::bind_front_handler(&CompleteWebSocketSession::on_read,
                                     shared_from_this()));
    }
    
    void on_read(beast::error_code ec, std::size_t) {
        if (ec) {
            if (ec == websocket::error::closed) {
                std::cout << "Connection closed normally\n";
            } else {
                std::cerr << "Read: " << ec.message() << "\n";
            }
            return;
        }
        
        ws_.text(ws_.got_text());
        ws_.async_write(buffer_.data(),
            beast::bind_front_handler(&CompleteWebSocketSession::on_write,
                                     shared_from_this()));
    }
    
    void on_write(beast::error_code ec, std::size_t) {
        if (ec) {
            std::cerr << "Write: " << ec.message() << "\n";
            return;
        }
        
        buffer_.consume(buffer_.size());
        do_read();
    }
};

The echo session shows the essential rule of Beast’s stream: at most one read and one write outstanding at a time. Here it is enforced by construction, because the next read is started only after the write completes. ws_.text(ws_.got_text()) preserves the message type, so a binary frame is echoed as binary. The buffer is consumed only in on_write, since async_write references the buffer’s memory until it completes; consuming earlier would send garbage. Setting timeout::suggested(role_type::server) before async_accept gives the handshake a time limit and enables keep-alive pings, which protects the server from clients that connect and never send anything.

Common errors

Error 1: Handshake 400 Bad Request

Symptom: the client’s handshake fails with websocket::error::upgrade_declined (the server answered with something other than 101); on a Beast server, async_accept fails with a specific error such as websocket::error::no_sec_key or no_upgrade. Cause:

  • Missing or malformed Sec-WebSocket-Key
  • Wrong Upgrade token / casing
  • Missing Connection: Upgrade Fix:
// Beast generates correct headers
// manual implementations must follow RFC 6455
ws_.async_handshake(host, path,
    [](beast::error_code ec) {
        if (ec == websocket::error::upgrade_declined) {
            std::cerr << "Server refused the upgrade: check Upgrade, Connection, Sec-WebSocket-Key\n";
        }
    });

When the client and server are both correct, look at what sits between them. A reverse proxy must be configured to forward the upgrade; with nginx that means proxy_http_version 1.1; plus proxy_set_header Upgrade $http_upgrade; and proxy_set_header Connection "upgrade";, because Upgrade and Connection are hop-by-hop headers that are not forwarded by default.

Error 2: bad_version (426 Upgrade Required)

Symptom: server returns 426 (a Beast server reports websocket::error::bad_sec_version) Cause: Sec-WebSocket-Version ≠ 13 Fix: Beast defaults to 13; manual stacks must send 13.

Error 3: connection_reset / connection_aborted

Symptom: read/write fails mid-flight Cause:

  • NAT/firewall idle timeout
  • server restart
  • flaky network Fix: heartbeats + reconnect policy
ws_.async_read(buffer_,
    [self = shared_from_this()](beast::error_code ec, std::size_t) {
        if (ec) {
            if (ec == net::error::connection_reset ||
                ec == net::error::connection_aborted) {
                self->schedule_reconnect();
            }
            return;
        }
        // ...
    });

Not every error should trigger a reconnect. websocket::error::closed means the peer completed a proper close handshake, which is normal; net::error::operation_aborted usually means your own code cancelled the operation during shutdown, and reconnecting in that case fights your own shutdown logic. Reconnect on transport failures, and on closed only if the close code says the server is going away (1001) or restarting.

Error 4: frame too big / payload too large

Symptom: websocket::error::message_too_big on the reading side; Beast then closes the connection with code 1009 (message too big) Cause: payload larger than read_message_max Fix:

// server: cap incoming messages
ws_.read_message_max(1024 * 1024);  // 1MB
// set the same limit on clients
ws_.read_message_max(1024 * 1024);

Error 5: random WSS drops (multithreaded)

Symptom: random drops or TLS/protocol errors in browsers, only under load with several I/O threads Cause: concurrent access to the same stream Fix: serialize with a strand

// stream bound to a strand
auto strand = net::make_strand(ioc);
websocket::stream<beast::tcp_stream> ws_(strand);
// completion handlers of ws_ now run on that strand by default;
// bind_executor is only needed for handlers created elsewhere
ws_.async_read(buffer_, net::bind_executor(strand, [](beast::error_code ec, std::size_t) { ... }));

Constructing the stream on a strand serializes the completion handlers, but it does not protect you from initiating an operation from another thread. Code that calls session->send(msg) from a worker thread or a different connection’s handler must net::post(ws_.get_executor(), ...) so the call itself runs on the strand. Forgetting that is the usual way a multithreaded Beast server ends up with two writes in flight on one stream, and the result is exactly this kind of non-reproducible failure. Running with ThreadSanitizer (-fsanitize=thread) in a load test usually finds it quickly.

Error 6: Mask required (client → server)

Symptom: server rejects client frames Cause: RFC 6455 requires masking for client→server frames Fix: Beast masks automatically; manual stacks must set MASK.

Error 7: Invalid UTF-8 (text frames)

Symptom: websocket::error::bad_frame_payload on the receiving side, and the connection is closed with code 1007 Cause: invalid UTF-8 in a text frame. RFC 6455 requires text messages to be valid UTF-8, and Beast validates incoming text frames. A common source is sending a std::string containing Latin-1 or truncated multi-byte data (for example a string cut at a fixed byte length) with the stream in text mode. Fix:

// send as binary or validate UTF-8 first
ws_.binary(true);
ws_.async_write(net::buffer(data), ...);

Error 8: Double read (overlapping async_read)

Symptom: UB / crashes (debug builds of Beast often stop on an assertion in the stream’s internal “soft mutex”) Cause: second async_read before the first completes, or a second async_write while one is in flight Fix: chain reads only inside the completion handler

void on_read(beast::error_code ec, std::size_t) {
    if (ec) return;
    // handle ...
    do_read();  // schedule next read here only
}

Error 9: connect or handshake hangs forever

Typical causes: firewall, wrong host/port, flaky Wi‑Fi.

// ❌ async_connect without deadline → can hang forever
beast::get_lowest_layer(ws_).async_connect(results, ...);
// ✅ set deadline / timeout on tcp_stream
beast::get_lowest_layer(ws_).expires_after(std::chrono::seconds(10));
beast::get_lowest_layer(ws_).async_connect(results,
    [](beast::error_code ec) {
        if (ec == net::error::operation_aborted) {
            std::cerr << "Connection timeout\n";
        }
    });

Error 10: sessions never freed (timers holding shared_ptr)

Cause: timers capture shared_from_this() and never cancel.

// ❌ timer keeps session alive forever
ping_timer_.async_wait([self = shared_from_this()](...) {
    self->ws_.async_ping(...);
});
// ✅ use weak_ptr for heartbeat callbacks
auto weak = std::weak_ptr<Session>(shared_from_this());
ping_timer_.async_wait([weak](beast::error_code ec) {
    auto self = weak.lock();
    if (!self || ec) return;
    // ...
});

The issue is not a reference cycle in the usual sense but a self-renewing reference: every timer handler holds a shared_ptr, and every handler schedules the next timer, so the session’s reference count never reaches zero even after the socket has died. The symptom is a slow memory leak and a count of “active sessions” that only grows. Besides weak_ptr, the other standard fix is to cancel the timers explicitly when the read loop ends with an error; the pending wait then completes with operation_aborted, the handler returns without rescheduling, and the last reference goes away.

Best practices

Cap message size

// rough guidance
// chat JSON: 64 KiB
// JSON API: 1MB
// binary streaming: up to 10 MiB (watch memory)
ws_.read_message_max(64 * 1024);

Reconnect with exponential backoff

// attempt_ is a member; reset it to 0 in on_handshake() after a successful connect
void schedule_reconnect() {
    auto delay = std::chrono::seconds(std::min(1 << std::min(attempt_, 6), 60));
    ++attempt_;
    
    reconnect_timer_.expires_after(delay);
    reconnect_timer_.async_wait(
        [self = shared_from_this()](beast::error_code ec) {
            if (!ec) {
                self->connect(self->host_, self->port_, self->path_);
            }
        });
}

The backoff has to reset on a successful handshake, not when the reconnect attempt is started; otherwise a server that accepts TCP and then fails every handshake is retried every second forever. Add random jitter to the delay (for example up to 30% extra) in any client that runs in many copies. Without it, after a server restart all clients disconnect at the same moment and reconnect in synchronized waves at 1 s, 2 s, 4 s and so on, which can knock the server over again just as it comes back. I have seen more outages prolonged by synchronized reconnects than caused by the original failure.

Graceful Close

void close() {
    ws_.async_close(websocket::close_code::normal,
        [self = shared_from_this()](beast::error_code ec) {
            if (ec) {
                beast::get_lowest_layer(self->ws_).close();
            }
        });
}

async_close sends a close frame and then keeps reading until the peer’s close frame arrives, discarding any data messages in between, and only then shuts down the TCP connection. That is why a pending async_read elsewhere completes with websocket::error::closed during a graceful close. If the peer never answers, the close waits until the stream’s timeout fires, which is another reason to configure one.

Fan-out: rooms and topic subscriptions

class ChatServer {
    std::set<std::shared_ptr<WebSocketSession>> sessions_;
    std::mutex mutex_;
    
public:
    void join(std::shared_ptr<WebSocketSession> session) {
        std::lock_guard<std::mutex> lock(mutex_);
        sessions_.insert(session);
    }
    
    void leave(std::shared_ptr<WebSocketSession> session) {
        std::lock_guard<std::mutex> lock(mutex_);
        sessions_.erase(session);
    }
    
    void broadcast(const std::string& message) {
        std::lock_guard<std::mutex> lock(mutex_);
        
        for (auto& session : sessions_) {
            session->send(message);
        }
    }
};

The mutex protects the set, not the sessions. That is only safe if send does no I/O itself and merely posts the message to the session’s strand, where a per-session queue writes it; if send wrote directly, one slow client would hold the lock and block the whole room. leave must also be called from every exit path of a session (read error, write error, timeout), or the set keeps dead sessions alive through their shared_ptr and broadcasts go to sockets that no longer exist.

Topic-based dashboard

class DashboardServer {
    // many sessions per topic
    std::map<std::string, std::vector<std::shared_ptr<WebSocketSession>>> subscribers_;
    
public:
    void subscribe(const std::string& topic, std::shared_ptr<WebSocketSession> session) {
        subscribers_[topic].push_back(std::move(session));
    }
    
    void publish(const std::string& topic, const nlohmann::json& data) {
        auto it = subscribers_.find(topic);
        if (it != subscribers_.end()) {
            const std::string payload = data.dump();   // serialize once
            for (auto& s : it->second) s->send(payload);
        }
    }
    
    // push metrics every second (example)
    void pushMetrics() {
        nlohmann::json metrics = {
            {"cpu", getCpuUsage()},
            {"memory", getMemoryUsage()},
            {"requests", getRequestCount()}
        };
        
        publish("metrics", metrics);
    }
};

Broadcast backpressure

// ❌ bad: thousands of concurrent writes
for (auto& session : sessions_) {
    session->ws_.async_write(...);  // floods executor
}
// ✅ good: per-session queues
void broadcast(const std::string& msg) {
    for (auto& session : sessions_) {
        session->enqueue(msg);
    }
}
void enqueue(const std::string& msg) {
    bool was_empty = write_queue_.empty();
    write_queue_.push(msg);
    if (was_empty) do_write();
}
void do_write() {
    if (write_queue_.empty()) return;
    ws_.async_write(net::buffer(write_queue_.front()),
        [this](beast::error_code ec, std::size_t) {
            if (!ec) {
                write_queue_.pop();
                do_write();
            }
        });
}

The queue guarantees one write in flight per session, which Beast requires anyway, and it also isolates slow clients: a client on a bad mobile link only grows its own queue instead of stalling the broadcast loop. The sketch leaves out three things production code needs. enqueue must run on the session’s strand, so broadcast should do net::post(session->ws_.get_executor(), [session, msg] { session->enqueue(msg); }). The lambda in do_write should capture shared_from_this() rather than this, or a session destroyed mid-write leaves a dangling pointer. And the queue needs a limit: if a client’s queue exceeds some size, close it with a policy-violation code rather than letting memory grow without bound. For large fan-out, storing messages as std::shared_ptr<const std::string> avoids copying the same payload once per session.

Configure timeouts

websocket::stream_base::timeout opt{
    std::chrono::seconds(30),  // handshake timeout
    std::chrono::seconds(30),  // idle timeout
    false                     // keepalive pings
};
ws_.set_option(opt);

Read this configuration carefully, because it is a common self-inflicted disconnect: with a 30-second idle timeout and keep-alive pings turned off, Beast closes any connection that is silent for 30 seconds, including perfectly healthy chat clients that simply have nothing to say. Either turn keep_alive_pings on (then the idle timeout means “no response to our ping”) or set the idle timeout to websocket::stream_base::none() and rely on your own heartbeat. Note also that Beast’s websocket timeout replaces the tcp_stream timeout once the handshake has completed; call beast::get_lowest_layer(ws_).expires_never() before setting it, as the Beast examples do, so the two do not fight.

Logging

ws_.async_handshake(host, path,
    [host, path](beast::error_code ec) {
        if (ec) {
            spdlog::error("WebSocket handshake failed: {} {} {}",
                host, path, ec.message());
        }
    });

Production patterns

Pattern 1: connection limits

class WebSocketServer {
    std::atomic<int> connection_count_{0};
    static constexpr int max_connections_ = 10000;
    
    void do_accept() {
        acceptor_.async_accept(
            [this](beast::error_code ec, tcp::socket socket) {
                if (ec) return;
                
                if (connection_count_.load() >= max_connections_) {
                    socket.close();
                    spdlog::warn("Connection limit reached");
                } else {
                    connection_count_++;
                    std::make_shared<Session>(std::move(socket),
                        [this]() { connection_count_--; })->run();
                }
                do_accept();
            });
    }
};

Pattern 2: Graceful shutdown

void shutdown() {
    acceptor_.close();
    
    for (auto& session : sessions_) {
        session->ws_.async_close(websocket::close_code::going_away,
            [](beast::error_code) {});
    }
    
    work_guard_.reset();
    // Do not call ioc_.stop() here: it would abandon the close handshakes started above.
    // Let run() return once all sessions finish, with a deadline timer as a fallback.
}

The order matters. Closing the acceptor first stops new connections; sending close code 1001 (“going away”) tells well-behaved clients that the server is restarting, so they can reconnect elsewhere with backoff rather than treating it as an error. Calling ioc_.stop() immediately after starting the closes, as many examples do, abandons every pending operation, so clients see an abrupt 1006 instead of 1001. Each async_close must also be initiated on that session’s strand (via net::post) when shutdown runs on another thread.

Pattern 3: Metrics

struct WebSocketMetrics {
    std::atomic<uint64_t> connections_total{0};
    std::atomic<uint64_t> connections_active{0};
    std::atomic<uint64_t> messages_received{0};
    std::atomic<uint64_t> messages_sent{0};
    std::atomic<uint64_t> errors_handshake{0};
    std::atomic<uint64_t> errors_read{0};
};
// export to Prometheus/Grafana/etc.
void on_handshake(beast::error_code ec) {
    if (ec) {
        metrics_.errors_handshake++;
        return;
    }
    metrics_.connections_active++;
}

Pattern 4: Subprotocol negotiation

// client
ws_.set_option(websocket::stream_base::decorator(
    [](websocket::request_type& req) {
        req.set(beast::http::field::sec_websocket_protocol,
               "chat, json");
    }));
// server: pick one during accept
ws_.set_option(websocket::stream_base::decorator(
    [](websocket::response_type& res) {
        res.set(beast::http::field::sec_websocket_protocol, "chat");
    }));

The server side of this example is simplified in a way that breaks real clients: a server may only answer with one of the subprotocols the client offered, and must omit the header entirely if the client offered none. Browsers enforce this; if the response names a protocol the client did not request, the browser fails the connection. The correct approach is to read the upgrade request first (http::async_read into a request), inspect its Sec-WebSocket-Protocol list, and set the response header in the decorator only when there is a match.

Pattern 5: WSS (TLS)

using ssl_stream = boost::asio::ssl::stream<beast::tcp_stream>;
websocket::stream<ssl_stream> wss_(ssl_ctx, net::make_strand(ioc));
// TLS first, then WebSocket upgrade (set SNI with SSL_set_tlsext_host_name before this)
wss_.next_layer().async_handshake(ssl::stream_base::client,
    [this](beast::error_code ec) {
        if (!ec) {
            wss_.async_handshake(host_, path_, ...);
        }
    });

With TLS there are three layers: TCP connect, TLS handshake on next_layer(), then the WebSocket upgrade. The step people miss is Server Name Indication. Most servers behind a CDN or shared load balancer host many certificates on one IP and pick one based on the SNI hostname; without it, the TLS handshake fails with an error such as sslv3 alert handshake failure, even though the same URL works in a browser. Set it with SSL_set_tlsext_host_name(wss_.next_layer().native_handle(), host.c_str()) before the handshake, and enable certificate verification (ssl_ctx.set_verify_mode(ssl::verify_peer) with a hostname check), because the default context verifies nothing.

Pattern 6: load balancer and sticky sessions

Stateful features often need sticky sessions (same client → same backend) unless you externalize all fan-out (Redis, etc.). Strictly speaking, a WebSocket connection is already pinned to one backend for its whole life; stickiness matters for what happens around it, such as an HTTP request that must reach the node holding the user’s socket, or a reconnect that should land on the node with the user’s buffered state. ip_hash is a blunt tool for this, since many users behind one corporate NAT all hash to the same backend. The proxy_read_timeout 3600s line is what keeps nginx from closing idle WebSocket connections after its default 60 seconds; a heartbeat shorter than that timeout achieves the same without holding dead connections for an hour.

# Nginx example
upstream websocket_backend {
    ip_hash;  # pin client IP to one upstream
    server 10.0.1.1:8080;
    server 10.0.1.2:8080;
}
location /ws {
    proxy_pass http://websocket_backend;
    proxy_http_version 1.1;
    proxy_set_header Upgrade $http_upgrade;
    proxy_set_header Connection "Upgrade";
    proxy_read_timeout 3600s;
    proxy_send_timeout 3600s;
}

Pattern 7: horizontal scale with Redis Pub/Sub

Broadcast across nodes when connections are not all local:

// Node A publishes after handling a chat event
redis.publish("chat:room1", message);
// Nodes B/C subscribed to the channel forward to local room members
redis.subscribe("chat:room1", [this](const std::string& msg) {
    for (auto& session : room1_sessions_) {
        session->send(msg);
    }
});

Pattern 8: health checks

// For Kubernetes/ECS, expose a plain HTTP /health (200 OK) on a separate port
// WebSocket-specific probes often use synthetic Ping/Pong or TCP checks

Checklist

Handshake

  • Random 16-byte key, Base64
  • Sec-WebSocket-Accept SHA1+magic+Base64
  • Upgrade: websocket, Connection: Upgrade
  • Sec-WebSocket-Version: 13

Frames

  • Client→server frames masked
  • Validate UTF-8 for text frames
  • Configure read_message_max
  • Close frames include code + optional reason

Ping/Pong

  • Ping every 20–30s
  • Reconnect if Pong missing ~10s
  • Server answers Ping with Pong

Errors

  • connection_reset → reconnect
  • message_too_big → enforce caps
  • Log handshake failures; retry with backoff

Production

  • Use strands under concurrency
  • Exponential backoff reconnects
  • Enforce max concurrent sockets
  • Graceful shutdown
  • Metrics and structured logging

References


WebSocket details to get right

TopicDetail
HandshakeHTTP Upgrade + Sec-WebSocket-Key / Accept
FramesFIN, opcode, mask bit, payload
MaskingRequired client→server
Ping/Pong~30s heartbeat; reconnect if Pong missing ~10s
Errors400/426, connection_reset, message_too_big
Productionstrands, backpressure, backoff, metrics

FAQ

Handshake returns 400

Verify Sec-WebSocket-Key, Upgrade, and Connection. Beast sets these correctly when you use async_handshake / async_accept.

WSS connections drop randomly under load

Construct the stream on a strand (net::make_strand) and make sure every operation on it is initiated from that strand (net::post from other threads).

Broadcast slows the server

Use per-session write queues and sequential writes (backpressure) instead of fanning out thousands of concurrent async_write calls.