WebSocket in C++ with Boost.Beast: Handshake, Frames, Ping/Pong and Common Failures

Introduction: “We need bidirectional real-time messaging”

Why HTTP polling breaks down

// ❌ Problem: polling wastes work
while (true) {
    auto response = httpGet("/api/messages");
    if (response.has_new_messages()) {
        process(response.messages);
    }
    std::this_thread::sleep_for(std::chrono::seconds(1));
}
// Issues:
// - Mostly empty responses
// - Up to one period of latency
// - Load scales with clients × poll rate

What goes wrong in production:

  • Latency bounded by the poll interval
  • Wasted requests when nothing changed
  • Load ≈ concurrent users × poll frequency
  • Cost from traffic and CPU spent on empty replies

More scenarios

Scenario 2: polling melts the server

Fifty thousand users polling once per second is fifty thousand HTTP transactions per second. WebSocket keeps one socket per subscriber, which usually shrinks overhead dramatically.

Scenario 3: market data latency

A five-second poll means quotes can be five seconds stale—unacceptable for low-latency trading. Push over WebSocket delivers updates as they arrive.

Scenario 4: lost messages after reconnect

Mobile backgrounds drop TCP sessions. Without sequence numbers or offsets, you cannot resume mid-stream after reconnect.

Scenario 5: corporate proxies block upgrades

Some networks block Upgrade: websocket. Fallback strategies include long polling over HTTPS or WSS on port 443.

That last point is more important than it looks. A plain ws:// connection on port 80 passes through any proxy that inspects HTTP, and a proxy that does not understand the upgrade may strip the headers or buffer the response until it “completes”, which never happens. With wss://, the proxy only sees an encrypted TLS tunnel (usually via CONNECT) and cannot interfere with the upgrade. In practice, WSS on 443 is the only configuration that works reliably from corporate and hotel networks, which is why production deployments rarely offer anything else.

Before committing to WebSocket, it is also worth asking whether the traffic is really bidirectional. For one-way server-to-client updates, such as notifications or a live dashboard, Server-Sent Events over ordinary HTTP are simpler: they reconnect automatically, pass through proxies like any HTTP response, and work with HTTP/2 multiplexing. WebSocket earns its extra complexity when the client also sends frequent messages, as in chat, games or collaborative editing.

How WebSocket helps:

// ✅ WebSocket: server pushes when data exists
ws.async_read(buffer, [&](beast::error_code ec, std::size_t) {
    if (!ec) {
        process(buffer);
        ws.async_read(buffer, ...);
    }
});
// Benefits:
// - Event-driven delivery
// - One connection, bidirectional
// - No periodic polling loop

Goals:

  • Understand the WebSocket protocol (handshake, frames)
  • Use websocket::stream from Boost.Beast
  • Implement Ping/Pong heartbeats
  • Handle errors and reconnect
  • Compare WebSocket vs HTTP polling
  • Sketch production patterns (chat, dashboards) Prerequisites: Boost.Beast 1.70+ After reading you should be able to reason about frames in Wireshark, ship a Beast client/server, and tune timeouts for real networks.

Mental model

Think of sockets as addresses and async I/O as delivery routes: Asio schedules handlers (like workers) across threads; a strand keeps one connection’s work serialized.


Practical focus: the edge cases that textbooks skip (thread safety of one stream, masking, proxy and load balancer timeouts).

WebSocket protocol layout

Connection lifecycle

sequenceDiagram
    participant C as Client
    participant S as Server
    
    C->>S: HTTP GET /ws\nUpgrade: websocket\nSec-WebSocket-Key: xxx
    S->>C: HTTP 101 Switching Protocols\nSec-WebSocket-Accept: yyy
    
    Note over C,S: WebSocket established
    
    C->>S: WebSocket Frame (Text)
    S->>C: WebSocket Frame (Text)
    C->>S: Ping
    S->>C: Pong
    C->>S: Close
    S->>C: Close

Connection state machine

stateDiagram-v2
    [*] --> Connecting: TCP connect
    Connecting --> Open: 101 Switching Protocols
    Connecting --> [*]: 400/403 etc.
    Open --> Closing: Close frame received
    Open --> [*]: Abrupt drop
    Closing --> [*]: Close complete

HTTP vs WebSocket

AspectHTTPWebSocket
ConnectionOften per requestLong-lived
DirectionRequest/responseBidirectional
OverheadLarge HTTP headersSmall frame header (2–14 B)
TimelinessNeeds pollingPush
LoadHigh with pollingEvent-driven

The state machine hides one detail that matters for server design: once the connection is open, there is no request/response pairing at all. Either side may send a message at any time, and messages are not replies to anything unless your application protocol says so. That is why WebSocket applications almost always define an envelope on top, with a message type, often a request ID for correlation, and a sequence number if the client must be able to resume after reconnecting. Without that, the “lost messages after reconnect” scenario above has no solution.

Handshake

Client request

GET /chat HTTP/1.1
Host: example.com
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==
Sec-WebSocket-Version: 13

Important headers:

  • Upgrade: websocket — request protocol switch
  • Connection: Upgrade — HTTP upgrade hop
  • Sec-WebSocket-Key — random 16 bytes, Base64
  • Sec-WebSocket-Version: 13 — only standardized version

Server response

HTTP/1.1 101 Switching Protocols
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Accept: s3pPLMBiTxaQ9kYGzzhZRbK+xOo=

Raw bytes (first lines of the client request as sent on the wire):

47 45 54 20 2f 63 68 61 74 20 48 54 54 50 2f 31  GET /chat HTTP/1
2e 31 0d 0a 48 6f 73 74 3a 20 65 78 61 6d 70 6c  .1..Host: exampl
65 2e 63 6f 6d 0d 0a 55 70 67 72 61 64 65 3a 20  e.com..Upgrade:
77 65 62 73 6f 63 6b 65 74 0d 0a 43 6f 6e 6e 65  websocket..Conne
63 74 69 6f 6e 3a 20 55 70 67 72 61 64 65 0d 0a  ction: Upgrade..

Computing Sec-WebSocket-Accept:

#include <openssl/sha.h>
#include <boost/beast/core/detail/base64.hpp>
std::string computeAccept(const std::string& key) {
    // RFC 6455: key + "258EAFA5-E914-47DA-95CA-C5AB0DC85B11"
    std::string magic = "258EAFA5-E914-47DA-95CA-C5AB0DC85B11";
    std::string input = key + magic;
    
    // SHA-1
    unsigned char hash[SHA_DIGEST_LENGTH];
    SHA1(reinterpret_cast<const unsigned char*>(input.c_str()), 
         input.size(), hash);
    
    // Base64
    std::string result;
    result.resize(boost::beast::detail::base64::encoded_size(SHA_DIGEST_LENGTH));
    result.resize(boost::beast::detail::base64::encode(
        &result[0], hash, SHA_DIGEST_LENGTH));
    
    return result;
}

You will not call this function when using Beast, which computes and verifies the accept value inside async_handshake and async_accept; it is shown so the value in a packet capture makes sense. The magic GUID is fixed by RFC 6455, so the exchange proves only that the server speaks WebSocket, not who it is. Authentication has to come from elsewhere, typically a cookie or token checked during the upgrade request, or a first message after the connection opens. Browsers cannot set custom headers such as Authorization on a WebSocket handshake, which is why tokens often travel in a cookie or a query parameter; the latter ends up in proxy access logs, so prefer short-lived tokens there.

Handshake flow

flowchart TB
    Start[TCP connect] --> ClientReq[Client: HTTP Upgrade]
    ClientReq --> ServerCheck{Server: valid?}
    ServerCheck -->|Yes| ServerResp[Server: 101 Switching Protocols]
    ServerCheck -->|No| ServerErr[Server: 400 Bad Request]
    ServerResp --> WSOpen[WebSocket open]
    ServerErr --> Close[Connection closed]
    WSOpen --> DataExchange[Data frames]
    
    style WSOpen fill:#4caf50
    style ServerErr fill:#f44336

Frames

Frame layout

graph LR
    A["FIN<br/>1 bit"] --> B["RSV<br/>3 bits"]
    B --> C["Opcode<br/>4 bits"]
    C --> D["Mask<br/>1 bit"]
    D --> E["Payload Len<br/>7 bits"]
    E --> F["Extended<br/>Payload Len"]
    F --> G["Masking Key<br/>4 bytes"]
    G --> H[Payload Data]
    
    style A fill:#ff9800
    style C fill:#4caf50
    style D fill:#2196f3

Opcodes

OpcodeValueRole
Continuation0x0Continues previous fragment
Text0x1UTF-8 text
Binary0x2Binary payload
Close0x8Close connection
Ping0x9Heartbeat probe
Pong0xAHeartbeat reply

Close frames

// Beast: graceful close
ws_.close(websocket::close_code::normal);
// Optional close reason string
ws_.close(websocket::close_code::normal, "server shutting down");
// In read handler when closed
if (ec == websocket::error::closed) {
    auto reason = ws_.reason();
    // reason.code (1000 normal, 1001 going away, ...)
    // reason.reason (UTF-8 string)
}

A WebSocket close is a two-way handshake, like TCP’s FIN exchange but one layer up: one side sends a Close frame, the other answers with its own Close frame, and only then is the TCP connection shut down, by the server. websocket::error::closed from a read means that exchange completed normally, so it should not be logged as an error. The synchronous ws_.close() shown here blocks until the peer answers; inside an asynchronous server, always use async_close, or one unresponsive client can stall a thread.

Masking

Rule: client → server frames must be masked.

// XOR mask
void mask_payload(uint8_t* data, size_t len, const uint8_t mask_key[4]) {
    for (size_t i = 0; i < len; ++i) {
        data[i] ^= mask_key[i % 4];
    }
}

Why: mitigates cache poisoning on misbehaving intermediaries.

Worked example: text “Hi”

Raw bytes when the client sends two bytes of UTF-8 text:

Byte 0: 0x81  (FIN=1, RSV=0, opcode Text)
Byte 1: 0x82  (MASK=1, payload len=2)
Bytes 2–5: 4-byte masking key (example 0x37 0xfa 0x21 0x3d)
Bytes 6–7: "Hi" XOR mask → 0x7b 0x9b (masked payload)

Sending with Beast (masking applied automatically in client role):

// Beast masks client→server frames for you
ws_.text(true);
ws_.async_write(net::buffer("Hi"),
    [](beast::error_code ec, std::size_t) {
        if (!ec) {
            // 2 B payload + 2–14 B header on the wire
        }
    });

Frames and messages are different things, and Beast mostly hides the difference. A message may be split into several frames (the first with the real opcode and FIN=0, the rest as Continuation frames, the last with FIN=1), and control frames may be interleaved between those fragments. async_read reassembles a complete message into the buffer, while async_read_some exposes partial data for streaming very large messages. Two small traps in the example above: net::buffer("Hi") on a string literal would include the terminating '\0' if you wrote sizeof-based code by hand, and a buffer passed to async_write must stay alive until the handler runs, which a literal does but a local std::string does not.

Beast WebSocket client

Minimal synchronous client

#include <boost/beast.hpp>
#include <boost/asio.hpp>
#include <iostream>
namespace beast = boost::beast;
namespace websocket = beast::websocket;
namespace net = boost::asio;
using tcp = net::ip::tcp;
class WebSocketClient {
    net::io_context& ioc_;
    websocket::stream<beast::tcp_stream> ws_;
    beast::flat_buffer buffer_;
    
public:
    explicit WebSocketClient(net::io_context& ioc)
        : ioc_(ioc), ws_(net::make_strand(ioc)) {}
    
    void connect(const std::string& host, const std::string& port) {
        // Resolve host
        tcp::resolver resolver(ioc_);
        auto results = resolver.resolve(host, port);
        
        // TCP connect (tcp_stream::connect tries each resolved endpoint)
        beast::get_lowest_layer(ws_).connect(results);
        
        // WebSocket handshake
        ws_.handshake(host, "/");
        
        std::cout << "WebSocket connected to " << host << "\n";
    }
    
    void send(const std::string& message) {
        ws_.write(net::buffer(message));
    }
    
    std::string receive() {
        buffer_.clear();
        ws_.read(buffer_);
        
        return beast::buffers_to_string(buffer_.data());
    }
    
    void close() {
        ws_.close(websocket::close_code::normal);
    }
};
int main() {
    net::io_context ioc;
    
    WebSocketClient client(ioc);
    client.connect("echo.websocket.org", "80");
    
    client.send("Hello, WebSocket!");
    std::string response = client.receive();
    std::cout << "Received: " << response << "\n";
    
    client.close();
}

The synchronous client is the easiest way to test a server, and each call maps directly to one protocol step: resolve, TCP connect, HTTP upgrade, then framed reads and writes. Every call blocks and throws boost::system::system_error on failure, so wrap main in a try block in real use. Public echo services have come and gone over the years, so if the host in this example does not answer, run the echo server from the next section locally and connect to localhost:8080. Note that the host passed to handshake becomes the HTTP Host header; for a server on a non-default port it should include the port ("localhost:8080"), and some servers reject the upgrade if it does not match what they expect.

The synchronous API is fine for tools and tests, but a server cannot dedicate a blocked thread to each of thousands of mostly idle connections, which is what the asynchronous version below is for.

Async client

class AsyncWebSocketClient : public std::enable_shared_from_this<AsyncWebSocketClient> {
    websocket::stream<beast::tcp_stream> ws_;
    beast::flat_buffer buffer_;
    
public:
    explicit AsyncWebSocketClient(net::io_context& ioc)
        : ws_(net::make_strand(ioc)) {}
    
    void connect(const std::string& host, const std::string& port) {
        tcp::resolver resolver(ws_.get_executor());
        
        resolver.async_resolve(host, port,
            [self = shared_from_this(), host](
                beast::error_code ec,
                tcp::resolver::results_type results) {
                
                if (ec) {
                    std::cerr << "Resolve error: " << ec.message() << "\n";
                    return;
                }
                
                // TCP connect
                beast::get_lowest_layer(self->ws_).async_connect(results,
                    [self, host](beast::error_code ec, tcp::endpoint) {
                        if (ec) {
                            std::cerr << "Connect error: " << ec.message() << "\n";
                            return;
                        }
                        
                        // WebSocket handshake
                        self->ws_.async_handshake(host, "/",
                            [self](beast::error_code ec) {
                                if (ec) {
                                    std::cerr << "Handshake error: " << ec.message() << "\n";
                                    return;
                                }
                                
                                std::cout << "WebSocket connected\n";
                                self->do_read();
                            });
                    });
            });
    }
    
    void send(const std::string& message) {
        ws_.async_write(net::buffer(message),
            [](beast::error_code ec, std::size_t) {
                if (ec) {
                    std::cerr << "Write error: " << ec.message() << "\n";
                }
            });
    }
    
private:
    void do_read() {
        auto self = shared_from_this();
        
        ws_.async_read(buffer_,
            [self](beast::error_code ec, std::size_t bytes) {
                if (ec) {
                    if (ec != websocket::error::closed) {
                        std::cerr << "Read error: " << ec.message() << "\n";
                    }
                    return;
                }
                
                std::cout << "Received: " 
                          << beast::buffers_to_string(self->buffer_.data()) << "\n";
                
                self->buffer_.clear();
                self->do_read();  // wait for next message
            });
    }
};

This version has three bugs that are typical of first asynchronous clients, and they are worth recognizing because they compile cleanly and often pass a quick test. First, tcp::resolver resolver is a local variable; when connect() returns, the resolver is destroyed, and destroying an Asio resolver cancels its pending operation, so the handler may fire with operation_aborted. Make it a member. Second, send() passes net::buffer(message) where message is the caller’s string; if the caller’s string goes out of scope before the write completes, Beast sends freed memory. Copy the message into storage the object owns. Third, calling send() twice in quick succession starts two overlapping async_write operations, which Beast does not allow. The fix for both of the last two is the same: a per-connection outgoing queue, where send appends to the queue and only starts a write when none is in flight, as shown in the deep-dive article.

Beast WebSocket server

Echo server sketch

class WebSocketSession : public std::enable_shared_from_this<WebSocketSession> {
    websocket::stream<beast::tcp_stream> ws_;
    beast::flat_buffer buffer_;
    
public:
    explicit WebSocketSession(tcp::socket socket)
        : ws_(std::move(socket)) {}
    
    void run() {
        // Suggested timeouts
        ws_.set_option(websocket::stream_base::timeout::suggested(
            beast::role_type::server));
        
        ws_.set_option(websocket::stream_base::decorator(
            [](websocket::response_type& res) {
                res.set(beast::http::field::server, "Beast WebSocket Server");
            }));
        
        // Accept HTTP upgrade
        ws_.async_accept(
            [self = shared_from_this()](beast::error_code ec) {
                if (ec) {
                    std::cerr << "Accept error: " << ec.message() << "\n";
                    return;
                }
                
                self->do_read();
            });
    }
    
private:
    void do_read() {
        auto self = shared_from_this();
        
        ws_.async_read(buffer_,
            [self](beast::error_code ec, std::size_t) {
                if (ec) {
                    if (ec == websocket::error::closed) {
                        std::cout << "Connection closed\n";
                    } else {
                        std::cerr << "Read error: " << ec.message() << "\n";
                    }
                    return;
                }
                
                // Echo: preserve text/binary flag
                self->ws_.text(self->ws_.got_text());
                self->ws_.async_write(self->buffer_.data(),
                    [self](beast::error_code ec, std::size_t) {
                        if (ec) {
                            std::cerr << "Write error: " << ec.message() << "\n";
                            return;
                        }
                        
                        self->buffer_.clear();
                        self->do_read();
                    });
            });
    }
};
class WebSocketServer {
    net::io_context& ioc_;
    tcp::acceptor acceptor_;
    
public:
    WebSocketServer(net::io_context& ioc, uint16_t port)
        : ioc_(ioc),
          acceptor_(ioc, tcp::endpoint(tcp::v4(), port)) {}
    
    void run() {
        do_accept();
    }
    
private:
    void do_accept() {
        acceptor_.async_accept(
            net::make_strand(ioc_),
            [this](beast::error_code ec, tcp::socket socket) {
                if (!ec) {
                    std::make_shared<WebSocketSession>(std::move(socket))->run();
                }
                
                do_accept();
            });
    }
};
int main() {
    net::io_context ioc{1};
    
    WebSocketServer server(ioc, 8080);
    server.run();
    
    std::cout << "WebSocket server listening on port 8080\n";
    
    ioc.run();
}

The acceptor hands each new socket a fresh strand (net::make_strand(ioc_)), and the session’s stream inherits that executor. With io_context ioc{1} and one thread calling run(), strands make no difference yet, but the structure is ready for more threads: each connection’s handlers stay serialized while different connections run in parallel. timeout::suggested(role_type::server) sets a handshake timeout and an idle timeout with automatic keep-alive pings, so a client that connects and then disappears is eventually cleaned up instead of holding a socket forever. The decorator only adds a Server header to the 101 response; it is also where you would add a negotiated subprotocol. Because do_accept re-arms itself even after an error, a transient failure such as running out of file descriptors (too many open files) produces a tight loop of failing accepts; production servers log the error and wait briefly before accepting again.

Ping/Pong Heartbeat

Ping/Pong sequence

sequenceDiagram
    participant C as Client
    participant S as Server
    
    Note over C: Ping every 30s
    C->>S: Ping
    S->>C: Pong
    
    Note over C: keepalive OK
    
    C->>S: Ping
    Note over S: no Pong (dead peer)
    
    Note over C: timeout → reconnect

Ping timer

class WebSocketClientWithPing : public std::enable_shared_from_this<WebSocketClientWithPing> {
    websocket::stream<beast::tcp_stream> ws_;
    beast::flat_buffer buffer_;
    net::steady_timer ping_timer_;
    
public:
    explicit WebSocketClientWithPing(net::io_context& ioc)
        : ws_(net::make_strand(ioc)),
          ping_timer_(ws_.get_executor()) {}
    
    void start_ping() {
        ping_timer_.expires_after(std::chrono::seconds(30));
        
        ping_timer_.async_wait(
            [self = shared_from_this()](beast::error_code ec) {
                if (ec) return;
                
                // Send ping
                self->ws_.async_ping({},
                    [self](beast::error_code ec) {
                        if (ec) {
                            std::cerr << "Ping error: " << ec.message() << "\n";
                            return;
                        }
                        
                        self->start_ping();  // schedule next ping
                    });
            });
    }
};

Automatic Pong

Beast replies to Ping with Pong by default. For logging or custom behavior:

ws_.control_callback(
    [](websocket::frame_type kind, beast::string_view) {
        if (kind == websocket::frame_type::ping) {
            std::cout << "Received Ping\n";
            // Beast sends the Pong itself; this callback is only a notification
        } else if (kind == websocket::frame_type::pong) {
            std::cout << "Received Pong\n";
        }
    });

Heartbeat with Pong timeout

Example pattern: reconnect if Pong is missing:

class WebSocketWithHeartbeat : public std::enable_shared_from_this<WebSocketWithHeartbeat> {
    websocket::stream<beast::tcp_stream> ws_;
    beast::flat_buffer buffer_;
    net::steady_timer ping_timer_;
    net::steady_timer pong_timer_;
    bool pong_received_ = true;
    
public:
    void start_heartbeat() {
        pong_received_ = true;
        schedule_ping();
    }
    
private:
    void schedule_ping() {
        ping_timer_.expires_after(std::chrono::seconds(30));
        ping_timer_.async_wait(
            [self = shared_from_this()](beast::error_code ec) {
                if (ec) return;
                
                if (!self->pong_received_) {
                    std::cerr << "Pong timeout - reconnecting\n";
                    self->reconnect();
                    return;
                }
                
                self->pong_received_ = false;
                self->ws_.async_ping({},
                    [self](beast::error_code ec) {
                        if (ec) return;
                        self->schedule_pong_timeout();
                        self->schedule_ping();
                    });
            });
    }
    
    void schedule_pong_timeout() {
        pong_timer_.expires_after(std::chrono::seconds(10));
        pong_timer_.async_wait(
            [self = shared_from_this()](beast::error_code ec) {
                if (ec) return;
                if (!self->pong_received_) {
                    self->ws_.close(websocket::close_code::normal);
                }
            });
    }
    
    void setup_control_callback() {
        ws_.control_callback(
            [self = shared_from_this()](websocket::frame_type kind, beast::string_view) {
                if (kind == websocket::frame_type::pong) {
                    self->pong_received_ = true;
                }
            });
    }
    
    void reconnect() { /* ... */ }
};

Two details make or break this pattern. The control callback only runs while an async_read is pending, because Beast processes incoming control frames as part of reading; a client that stops reading will never see a Pong and will “time out” a healthy connection. And the timeout handler should use async_close rather than the blocking close shown here, which would stall the strand while waiting for a peer that is, by assumption, not answering. If all you need is dead-peer detection, Beast can do this for you: set the websocket timeout option with keep_alive_pings = true and an idle_timeout, and Beast pings at half the idle interval and fails the connection when nothing arrives. The hand-written version is mainly useful when you want to measure round-trip time or apply your own policy.

Before you put this in production

The examples above are enough to talk WebSocket end to end, but a service that stays up needs more: strict handshake validation and Origin checks, a message size cap, timeouts on connect and handshake, one strand per connection so reads and writes never overlap, timers that do not keep dead sessions alive, reconnect with backoff and jitter on the client, and backpressure when one slow client cannot keep up with a broadcast. Each of these, with the error it prevents, is covered in WebSocket in production (#30-2).


Performance snapshot

WebSocket vs HTTP polling

Example: 10k concurrent users, one event per user per minute.

ModeRequests/minBandwidth (illustrative)Latency
HTTP poll (1s)600,000largeup to 1s
Long poll~10,000mediumup to hold window
WebSocket~10,000smallpush

Takeaway: polling wastes bandwidth when updates are rare; WebSocket removes periodic HTTP overhead.

Order-of-magnitude resources

Concurrent socketsRAMCPU
1,000tens of MBlow %
10,000hundreds of MBmid teens %
100,000multi‑GBworkload dependent

These numbers are orders of magnitude, not measurements, and the dominant factor is usually buffers, not the socket itself. Each connection carries a read buffer, a queue of outgoing messages, TLS state (tens of kilobytes per connection with OpenSSL), and kernel socket buffers. A flat buffer that grew to 1 MiB for one large message keeps that capacity unless you shrink it. At high connection counts, operating-system limits bite before CPU does: the per-process file descriptor limit (ulimit -n, often 1024 by default) causes accept failures with too many open files long before memory runs out.

Scaling beyond one process

Room fan-out, sticky sessions behind a load balancer, cross-node broadcast through Redis Pub/Sub and health checks are covered in #30-2, next to the broadcast backpressure they depend on.


References


What a WebSocket server has to get right

TopicDetail
Wire formatHTTP upgrade → framed messages
HandshakeSec-WebSocket-Key / Sec-WebSocket-Accept
FramesOpcodes for text, binary, ping, pong, close
MaskingRequired client→server
HeartbeatPing/Pong + app-level keepalives
OpsWSS, LB timeouts, Redis fan-out, metrics

FAQ

How much better than polling?

Polling issues scale with users × poll rate; WebSocket removes that multiplier for push-heavy workloads.

My WSS connections drop randomly under load

Serialize every async WebSocket operation on the stream’s own strand (see above), and check load balancer idle timeouts against your heartbeat interval.