WebSocket in C++ with Boost.Beast: Handshake, Frames, Ping/Pong and Common Failures
Introduction: “We need bidirectional real-time messaging”
Why HTTP polling breaks down
// ❌ Problem: polling wastes work
while (true) {
auto response = httpGet("/api/messages");
if (response.has_new_messages()) {
process(response.messages);
}
std::this_thread::sleep_for(std::chrono::seconds(1));
}
// Issues:
// - Mostly empty responses
// - Up to one period of latency
// - Load scales with clients × poll rate
What goes wrong in production:
- Latency bounded by the poll interval
- Wasted requests when nothing changed
- Load ≈ concurrent users × poll frequency
- Cost from traffic and CPU spent on empty replies
More scenarios
Scenario 2: polling melts the server
Fifty thousand users polling once per second is fifty thousand HTTP transactions per second. WebSocket keeps one socket per subscriber, which usually shrinks overhead dramatically.
Scenario 3: market data latency
A five-second poll means quotes can be five seconds stale—unacceptable for low-latency trading. Push over WebSocket delivers updates as they arrive.
Scenario 4: lost messages after reconnect
Mobile backgrounds drop TCP sessions. Without sequence numbers or offsets, you cannot resume mid-stream after reconnect.
Scenario 5: corporate proxies block upgrades
Some networks block Upgrade: websocket. Fallback strategies include long polling over HTTPS or WSS on port 443.
That last point is more important than it looks. A plain ws:// connection on port 80 passes through any proxy that inspects HTTP, and a proxy that does not understand the upgrade may strip the headers or buffer the response until it “completes”, which never happens. With wss://, the proxy only sees an encrypted TLS tunnel (usually via CONNECT) and cannot interfere with the upgrade. In practice, WSS on 443 is the only configuration that works reliably from corporate and hotel networks, which is why production deployments rarely offer anything else.
Before committing to WebSocket, it is also worth asking whether the traffic is really bidirectional. For one-way server-to-client updates, such as notifications or a live dashboard, Server-Sent Events over ordinary HTTP are simpler: they reconnect automatically, pass through proxies like any HTTP response, and work with HTTP/2 multiplexing. WebSocket earns its extra complexity when the client also sends frequent messages, as in chat, games or collaborative editing.
How WebSocket helps:
// ✅ WebSocket: server pushes when data exists
ws.async_read(buffer, [&](beast::error_code ec, std::size_t) {
if (!ec) {
process(buffer);
ws.async_read(buffer, ...);
}
});
// Benefits:
// - Event-driven delivery
// - One connection, bidirectional
// - No periodic polling loop
Goals:
- Understand the WebSocket protocol (handshake, frames)
- Use
websocket::streamfrom Boost.Beast - Implement Ping/Pong heartbeats
- Handle errors and reconnect
- Compare WebSocket vs HTTP polling
- Sketch production patterns (chat, dashboards) Prerequisites: Boost.Beast 1.70+ After reading you should be able to reason about frames in Wireshark, ship a Beast client/server, and tune timeouts for real networks.
Mental model
Think of sockets as addresses and async I/O as delivery routes: Asio schedules handlers (like workers) across threads; a strand keeps one connection’s work serialized.
Practical focus: the edge cases that textbooks skip (thread safety of one stream, masking, proxy and load balancer timeouts).
WebSocket protocol layout
Connection lifecycle
sequenceDiagram
participant C as Client
participant S as Server
C->>S: HTTP GET /ws\nUpgrade: websocket\nSec-WebSocket-Key: xxx
S->>C: HTTP 101 Switching Protocols\nSec-WebSocket-Accept: yyy
Note over C,S: WebSocket established
C->>S: WebSocket Frame (Text)
S->>C: WebSocket Frame (Text)
C->>S: Ping
S->>C: Pong
C->>S: Close
S->>C: Close
Connection state machine
stateDiagram-v2
[*] --> Connecting: TCP connect
Connecting --> Open: 101 Switching Protocols
Connecting --> [*]: 400/403 etc.
Open --> Closing: Close frame received
Open --> [*]: Abrupt drop
Closing --> [*]: Close complete
HTTP vs WebSocket
| Aspect | HTTP | WebSocket |
|---|---|---|
| Connection | Often per request | Long-lived |
| Direction | Request/response | Bidirectional |
| Overhead | Large HTTP headers | Small frame header (2–14 B) |
| Timeliness | Needs polling | Push |
| Load | High with polling | Event-driven |
The state machine hides one detail that matters for server design: once the connection is open, there is no request/response pairing at all. Either side may send a message at any time, and messages are not replies to anything unless your application protocol says so. That is why WebSocket applications almost always define an envelope on top, with a message type, often a request ID for correlation, and a sequence number if the client must be able to resume after reconnecting. Without that, the “lost messages after reconnect” scenario above has no solution.
Handshake
Client request
GET /chat HTTP/1.1
Host: example.com
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==
Sec-WebSocket-Version: 13
Important headers:
Upgrade: websocket— request protocol switchConnection: Upgrade— HTTP upgrade hopSec-WebSocket-Key— random 16 bytes, Base64Sec-WebSocket-Version: 13— only standardized version
Server response
HTTP/1.1 101 Switching Protocols
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Accept: s3pPLMBiTxaQ9kYGzzhZRbK+xOo=
Raw bytes (first lines of the client request as sent on the wire):
47 45 54 20 2f 63 68 61 74 20 48 54 54 50 2f 31 GET /chat HTTP/1
2e 31 0d 0a 48 6f 73 74 3a 20 65 78 61 6d 70 6c .1..Host: exampl
65 2e 63 6f 6d 0d 0a 55 70 67 72 61 64 65 3a 20 e.com..Upgrade:
77 65 62 73 6f 63 6b 65 74 0d 0a 43 6f 6e 6e 65 websocket..Conne
63 74 69 6f 6e 3a 20 55 70 67 72 61 64 65 0d 0a ction: Upgrade..
Computing Sec-WebSocket-Accept:
#include <openssl/sha.h>
#include <boost/beast/core/detail/base64.hpp>
std::string computeAccept(const std::string& key) {
// RFC 6455: key + "258EAFA5-E914-47DA-95CA-C5AB0DC85B11"
std::string magic = "258EAFA5-E914-47DA-95CA-C5AB0DC85B11";
std::string input = key + magic;
// SHA-1
unsigned char hash[SHA_DIGEST_LENGTH];
SHA1(reinterpret_cast<const unsigned char*>(input.c_str()),
input.size(), hash);
// Base64
std::string result;
result.resize(boost::beast::detail::base64::encoded_size(SHA_DIGEST_LENGTH));
result.resize(boost::beast::detail::base64::encode(
&result[0], hash, SHA_DIGEST_LENGTH));
return result;
}
You will not call this function when using Beast, which computes and verifies the accept value inside async_handshake and async_accept; it is shown so the value in a packet capture makes sense. The magic GUID is fixed by RFC 6455, so the exchange proves only that the server speaks WebSocket, not who it is. Authentication has to come from elsewhere, typically a cookie or token checked during the upgrade request, or a first message after the connection opens. Browsers cannot set custom headers such as Authorization on a WebSocket handshake, which is why tokens often travel in a cookie or a query parameter; the latter ends up in proxy access logs, so prefer short-lived tokens there.
Handshake flow
flowchart TB
Start[TCP connect] --> ClientReq[Client: HTTP Upgrade]
ClientReq --> ServerCheck{Server: valid?}
ServerCheck -->|Yes| ServerResp[Server: 101 Switching Protocols]
ServerCheck -->|No| ServerErr[Server: 400 Bad Request]
ServerResp --> WSOpen[WebSocket open]
ServerErr --> Close[Connection closed]
WSOpen --> DataExchange[Data frames]
style WSOpen fill:#4caf50
style ServerErr fill:#f44336
Frames
Frame layout
graph LR
A["FIN<br/>1 bit"] --> B["RSV<br/>3 bits"]
B --> C["Opcode<br/>4 bits"]
C --> D["Mask<br/>1 bit"]
D --> E["Payload Len<br/>7 bits"]
E --> F["Extended<br/>Payload Len"]
F --> G["Masking Key<br/>4 bytes"]
G --> H[Payload Data]
style A fill:#ff9800
style C fill:#4caf50
style D fill:#2196f3
Opcodes
| Opcode | Value | Role |
|---|---|---|
| Continuation | 0x0 | Continues previous fragment |
| Text | 0x1 | UTF-8 text |
| Binary | 0x2 | Binary payload |
| Close | 0x8 | Close connection |
| Ping | 0x9 | Heartbeat probe |
| Pong | 0xA | Heartbeat reply |
Close frames
// Beast: graceful close
ws_.close(websocket::close_code::normal);
// Optional close reason string
ws_.close(websocket::close_code::normal, "server shutting down");
// In read handler when closed
if (ec == websocket::error::closed) {
auto reason = ws_.reason();
// reason.code (1000 normal, 1001 going away, ...)
// reason.reason (UTF-8 string)
}
A WebSocket close is a two-way handshake, like TCP’s FIN exchange but one layer up: one side sends a Close frame, the other answers with its own Close frame, and only then is the TCP connection shut down, by the server. websocket::error::closed from a read means that exchange completed normally, so it should not be logged as an error. The synchronous ws_.close() shown here blocks until the peer answers; inside an asynchronous server, always use async_close, or one unresponsive client can stall a thread.
Masking
Rule: client → server frames must be masked.
// XOR mask
void mask_payload(uint8_t* data, size_t len, const uint8_t mask_key[4]) {
for (size_t i = 0; i < len; ++i) {
data[i] ^= mask_key[i % 4];
}
}
Why: mitigates cache poisoning on misbehaving intermediaries.
Worked example: text “Hi”
Raw bytes when the client sends two bytes of UTF-8 text:
Byte 0: 0x81 (FIN=1, RSV=0, opcode Text)
Byte 1: 0x82 (MASK=1, payload len=2)
Bytes 2–5: 4-byte masking key (example 0x37 0xfa 0x21 0x3d)
Bytes 6–7: "Hi" XOR mask → 0x7b 0x9b (masked payload)
Sending with Beast (masking applied automatically in client role):
// Beast masks client→server frames for you
ws_.text(true);
ws_.async_write(net::buffer("Hi"),
[](beast::error_code ec, std::size_t) {
if (!ec) {
// 2 B payload + 2–14 B header on the wire
}
});
Frames and messages are different things, and Beast mostly hides the difference. A message may be split into several frames (the first with the real opcode and FIN=0, the rest as Continuation frames, the last with FIN=1), and control frames may be interleaved between those fragments. async_read reassembles a complete message into the buffer, while async_read_some exposes partial data for streaming very large messages. Two small traps in the example above: net::buffer("Hi") on a string literal would include the terminating '\0' if you wrote sizeof-based code by hand, and a buffer passed to async_write must stay alive until the handler runs, which a literal does but a local std::string does not.
Beast WebSocket client
Minimal synchronous client
#include <boost/beast.hpp>
#include <boost/asio.hpp>
#include <iostream>
namespace beast = boost::beast;
namespace websocket = beast::websocket;
namespace net = boost::asio;
using tcp = net::ip::tcp;
class WebSocketClient {
net::io_context& ioc_;
websocket::stream<beast::tcp_stream> ws_;
beast::flat_buffer buffer_;
public:
explicit WebSocketClient(net::io_context& ioc)
: ioc_(ioc), ws_(net::make_strand(ioc)) {}
void connect(const std::string& host, const std::string& port) {
// Resolve host
tcp::resolver resolver(ioc_);
auto results = resolver.resolve(host, port);
// TCP connect (tcp_stream::connect tries each resolved endpoint)
beast::get_lowest_layer(ws_).connect(results);
// WebSocket handshake
ws_.handshake(host, "/");
std::cout << "WebSocket connected to " << host << "\n";
}
void send(const std::string& message) {
ws_.write(net::buffer(message));
}
std::string receive() {
buffer_.clear();
ws_.read(buffer_);
return beast::buffers_to_string(buffer_.data());
}
void close() {
ws_.close(websocket::close_code::normal);
}
};
int main() {
net::io_context ioc;
WebSocketClient client(ioc);
client.connect("echo.websocket.org", "80");
client.send("Hello, WebSocket!");
std::string response = client.receive();
std::cout << "Received: " << response << "\n";
client.close();
}
The synchronous client is the easiest way to test a server, and each call maps directly to one protocol step: resolve, TCP connect, HTTP upgrade, then framed reads and writes. Every call blocks and throws boost::system::system_error on failure, so wrap main in a try block in real use. Public echo services have come and gone over the years, so if the host in this example does not answer, run the echo server from the next section locally and connect to localhost:8080. Note that the host passed to handshake becomes the HTTP Host header; for a server on a non-default port it should include the port ("localhost:8080"), and some servers reject the upgrade if it does not match what they expect.
The synchronous API is fine for tools and tests, but a server cannot dedicate a blocked thread to each of thousands of mostly idle connections, which is what the asynchronous version below is for.
Async client
class AsyncWebSocketClient : public std::enable_shared_from_this<AsyncWebSocketClient> {
websocket::stream<beast::tcp_stream> ws_;
beast::flat_buffer buffer_;
public:
explicit AsyncWebSocketClient(net::io_context& ioc)
: ws_(net::make_strand(ioc)) {}
void connect(const std::string& host, const std::string& port) {
tcp::resolver resolver(ws_.get_executor());
resolver.async_resolve(host, port,
[self = shared_from_this(), host](
beast::error_code ec,
tcp::resolver::results_type results) {
if (ec) {
std::cerr << "Resolve error: " << ec.message() << "\n";
return;
}
// TCP connect
beast::get_lowest_layer(self->ws_).async_connect(results,
[self, host](beast::error_code ec, tcp::endpoint) {
if (ec) {
std::cerr << "Connect error: " << ec.message() << "\n";
return;
}
// WebSocket handshake
self->ws_.async_handshake(host, "/",
[self](beast::error_code ec) {
if (ec) {
std::cerr << "Handshake error: " << ec.message() << "\n";
return;
}
std::cout << "WebSocket connected\n";
self->do_read();
});
});
});
}
void send(const std::string& message) {
ws_.async_write(net::buffer(message),
[](beast::error_code ec, std::size_t) {
if (ec) {
std::cerr << "Write error: " << ec.message() << "\n";
}
});
}
private:
void do_read() {
auto self = shared_from_this();
ws_.async_read(buffer_,
[self](beast::error_code ec, std::size_t bytes) {
if (ec) {
if (ec != websocket::error::closed) {
std::cerr << "Read error: " << ec.message() << "\n";
}
return;
}
std::cout << "Received: "
<< beast::buffers_to_string(self->buffer_.data()) << "\n";
self->buffer_.clear();
self->do_read(); // wait for next message
});
}
};
This version has three bugs that are typical of first asynchronous clients, and they are worth recognizing because they compile cleanly and often pass a quick test. First, tcp::resolver resolver is a local variable; when connect() returns, the resolver is destroyed, and destroying an Asio resolver cancels its pending operation, so the handler may fire with operation_aborted. Make it a member. Second, send() passes net::buffer(message) where message is the caller’s string; if the caller’s string goes out of scope before the write completes, Beast sends freed memory. Copy the message into storage the object owns. Third, calling send() twice in quick succession starts two overlapping async_write operations, which Beast does not allow. The fix for both of the last two is the same: a per-connection outgoing queue, where send appends to the queue and only starts a write when none is in flight, as shown in the deep-dive article.
Beast WebSocket server
Echo server sketch
class WebSocketSession : public std::enable_shared_from_this<WebSocketSession> {
websocket::stream<beast::tcp_stream> ws_;
beast::flat_buffer buffer_;
public:
explicit WebSocketSession(tcp::socket socket)
: ws_(std::move(socket)) {}
void run() {
// Suggested timeouts
ws_.set_option(websocket::stream_base::timeout::suggested(
beast::role_type::server));
ws_.set_option(websocket::stream_base::decorator(
[](websocket::response_type& res) {
res.set(beast::http::field::server, "Beast WebSocket Server");
}));
// Accept HTTP upgrade
ws_.async_accept(
[self = shared_from_this()](beast::error_code ec) {
if (ec) {
std::cerr << "Accept error: " << ec.message() << "\n";
return;
}
self->do_read();
});
}
private:
void do_read() {
auto self = shared_from_this();
ws_.async_read(buffer_,
[self](beast::error_code ec, std::size_t) {
if (ec) {
if (ec == websocket::error::closed) {
std::cout << "Connection closed\n";
} else {
std::cerr << "Read error: " << ec.message() << "\n";
}
return;
}
// Echo: preserve text/binary flag
self->ws_.text(self->ws_.got_text());
self->ws_.async_write(self->buffer_.data(),
[self](beast::error_code ec, std::size_t) {
if (ec) {
std::cerr << "Write error: " << ec.message() << "\n";
return;
}
self->buffer_.clear();
self->do_read();
});
});
}
};
class WebSocketServer {
net::io_context& ioc_;
tcp::acceptor acceptor_;
public:
WebSocketServer(net::io_context& ioc, uint16_t port)
: ioc_(ioc),
acceptor_(ioc, tcp::endpoint(tcp::v4(), port)) {}
void run() {
do_accept();
}
private:
void do_accept() {
acceptor_.async_accept(
net::make_strand(ioc_),
[this](beast::error_code ec, tcp::socket socket) {
if (!ec) {
std::make_shared<WebSocketSession>(std::move(socket))->run();
}
do_accept();
});
}
};
int main() {
net::io_context ioc{1};
WebSocketServer server(ioc, 8080);
server.run();
std::cout << "WebSocket server listening on port 8080\n";
ioc.run();
}
The acceptor hands each new socket a fresh strand (net::make_strand(ioc_)), and the session’s stream inherits that executor. With io_context ioc{1} and one thread calling run(), strands make no difference yet, but the structure is ready for more threads: each connection’s handlers stay serialized while different connections run in parallel. timeout::suggested(role_type::server) sets a handshake timeout and an idle timeout with automatic keep-alive pings, so a client that connects and then disappears is eventually cleaned up instead of holding a socket forever. The decorator only adds a Server header to the 101 response; it is also where you would add a negotiated subprotocol. Because do_accept re-arms itself even after an error, a transient failure such as running out of file descriptors (too many open files) produces a tight loop of failing accepts; production servers log the error and wait briefly before accepting again.
Ping/Pong Heartbeat
Ping/Pong sequence
sequenceDiagram
participant C as Client
participant S as Server
Note over C: Ping every 30s
C->>S: Ping
S->>C: Pong
Note over C: keepalive OK
C->>S: Ping
Note over S: no Pong (dead peer)
Note over C: timeout → reconnect
Ping timer
class WebSocketClientWithPing : public std::enable_shared_from_this<WebSocketClientWithPing> {
websocket::stream<beast::tcp_stream> ws_;
beast::flat_buffer buffer_;
net::steady_timer ping_timer_;
public:
explicit WebSocketClientWithPing(net::io_context& ioc)
: ws_(net::make_strand(ioc)),
ping_timer_(ws_.get_executor()) {}
void start_ping() {
ping_timer_.expires_after(std::chrono::seconds(30));
ping_timer_.async_wait(
[self = shared_from_this()](beast::error_code ec) {
if (ec) return;
// Send ping
self->ws_.async_ping({},
[self](beast::error_code ec) {
if (ec) {
std::cerr << "Ping error: " << ec.message() << "\n";
return;
}
self->start_ping(); // schedule next ping
});
});
}
};
Automatic Pong
Beast replies to Ping with Pong by default. For logging or custom behavior:
ws_.control_callback(
[](websocket::frame_type kind, beast::string_view) {
if (kind == websocket::frame_type::ping) {
std::cout << "Received Ping\n";
// Beast sends the Pong itself; this callback is only a notification
} else if (kind == websocket::frame_type::pong) {
std::cout << "Received Pong\n";
}
});
Heartbeat with Pong timeout
Example pattern: reconnect if Pong is missing:
class WebSocketWithHeartbeat : public std::enable_shared_from_this<WebSocketWithHeartbeat> {
websocket::stream<beast::tcp_stream> ws_;
beast::flat_buffer buffer_;
net::steady_timer ping_timer_;
net::steady_timer pong_timer_;
bool pong_received_ = true;
public:
void start_heartbeat() {
pong_received_ = true;
schedule_ping();
}
private:
void schedule_ping() {
ping_timer_.expires_after(std::chrono::seconds(30));
ping_timer_.async_wait(
[self = shared_from_this()](beast::error_code ec) {
if (ec) return;
if (!self->pong_received_) {
std::cerr << "Pong timeout - reconnecting\n";
self->reconnect();
return;
}
self->pong_received_ = false;
self->ws_.async_ping({},
[self](beast::error_code ec) {
if (ec) return;
self->schedule_pong_timeout();
self->schedule_ping();
});
});
}
void schedule_pong_timeout() {
pong_timer_.expires_after(std::chrono::seconds(10));
pong_timer_.async_wait(
[self = shared_from_this()](beast::error_code ec) {
if (ec) return;
if (!self->pong_received_) {
self->ws_.close(websocket::close_code::normal);
}
});
}
void setup_control_callback() {
ws_.control_callback(
[self = shared_from_this()](websocket::frame_type kind, beast::string_view) {
if (kind == websocket::frame_type::pong) {
self->pong_received_ = true;
}
});
}
void reconnect() { /* ... */ }
};
Two details make or break this pattern. The control callback only runs while an async_read is pending, because Beast processes incoming control frames as part of reading; a client that stops reading will never see a Pong and will “time out” a healthy connection. And the timeout handler should use async_close rather than the blocking close shown here, which would stall the strand while waiting for a peer that is, by assumption, not answering. If all you need is dead-peer detection, Beast can do this for you: set the websocket timeout option with keep_alive_pings = true and an idle_timeout, and Beast pings at half the idle interval and fails the connection when nothing arrives. The hand-written version is mainly useful when you want to measure round-trip time or apply your own policy.
Before you put this in production
The examples above are enough to talk WebSocket end to end, but a service that stays up needs more: strict handshake validation and Origin checks, a message size cap, timeouts on connect and handshake, one strand per connection so reads and writes never overlap, timers that do not keep dead sessions alive, reconnect with backoff and jitter on the client, and backpressure when one slow client cannot keep up with a broadcast. Each of these, with the error it prevents, is covered in WebSocket in production (#30-2).
Performance snapshot
WebSocket vs HTTP polling
Example: 10k concurrent users, one event per user per minute.
| Mode | Requests/min | Bandwidth (illustrative) | Latency |
|---|---|---|---|
| HTTP poll (1s) | 600,000 | large | up to 1s |
| Long poll | ~10,000 | medium | up to hold window |
| WebSocket | ~10,000 | small | push |
Takeaway: polling wastes bandwidth when updates are rare; WebSocket removes periodic HTTP overhead.
Order-of-magnitude resources
| Concurrent sockets | RAM | CPU |
|---|---|---|
| 1,000 | tens of MB | low % |
| 10,000 | hundreds of MB | mid teens % |
| 100,000 | multi‑GB | workload dependent |
These numbers are orders of magnitude, not measurements, and the dominant factor is usually buffers, not the socket itself. Each connection carries a read buffer, a queue of outgoing messages, TLS state (tens of kilobytes per connection with OpenSSL), and kernel socket buffers. A flat buffer that grew to 1 MiB for one large message keeps that capacity unless you shrink it. At high connection counts, operating-system limits bite before CPU does: the per-process file descriptor limit (ulimit -n, often 1024 by default) causes accept failures with too many open files long before memory runs out.
Scaling beyond one process
Room fan-out, sticky sessions behind a load balancer, cross-node broadcast through Redis Pub/Sub and health checks are covered in #30-2, next to the broadcast backpressure they depend on.
References
What a WebSocket server has to get right
| Topic | Detail |
|---|---|
| Wire format | HTTP upgrade → framed messages |
| Handshake | Sec-WebSocket-Key / Sec-WebSocket-Accept |
| Frames | Opcodes for text, binary, ping, pong, close |
| Masking | Required client→server |
| Heartbeat | Ping/Pong + app-level keepalives |
| Ops | WSS, LB timeouts, Redis fan-out, metrics |
FAQ
How much better than polling?
Polling issues scale with users × poll rate; WebSocket removes that multiplier for push-heavy workloads.
My WSS connections drop randomly under load
Serialize every async WebSocket operation on the stream’s own strand (see above), and check load balancer idle timeouts against your heartbeat interval.