Real-Time Apps with WebSocket and Socket.io: RFC 6455 Frames, Rooms and Redis Scale-Out

Key takeaways

Hands-on guide to WebSocket and Socket.io: native API, rooms, Redis adapter, and the protocol layer—HTTP upgrade, frames, text vs binary frames, Ping/Pong heartbeats, and production patterns.

At a glance

This guide covers the browser WebSocket API, Socket.io (rooms, broadcast), a chat example, horizontal scaling with Redis, and RFC 6455 internals so you can debug proxies, handshakes, and heartbeats in production.

When you need real time

Polling wastes requests. Calling an API every second means most responses say “nothing new”, and each one pays for headers, authentication and a database query. A WebSocket keeps one connection open and sends only when something happens.

Polling adds delay. An update waits on average half the poll interval before the client sees it. A pushed message arrives as soon as the server sends it.

Collaboration needs both directions. Live cursors, co-editing and multiplayer state are streams of small messages from every participant, which map naturally to a persistent bidirectional channel.

The cost is that the server now holds state: every open connection is a socket, some memory and a place where messages can pile up. Load balancers, deploys and horizontal scaling all become harder than with stateless HTTP. If updates only flow from server to client, Server-Sent Events are often the simpler choice (see the FAQ).


What is WebSocket?

Core properties

WebSocket is a standardized protocol for full-duplex communication over a long-lived connection.

Benefits:

  • Bidirectional — server and client can both send at any time
  • Low latency — no new request per message, so delivery time is roughly the network round trip
  • Efficient — a frame header is 2–14 bytes, instead of repeated HTTP headers per message
  • Built into browsers — WebSocket API

HTTP vs WebSocket:

  • HTTP — request/response, often short-lived connections
  • WebSocket — persistent session after the upgrade

Protocol internals: upgrade, frames, text/binary, heartbeats (RFC 6455)

Libraries (ws, Socket.io, browser WebSocket) hide framing and masking. For production incidents—proxies closing idle connections, 400 on handshake, or invalid UTF-8 closes—you need the RFC 6455 skeleton.

From HTTP to WebSocket: the upgrade handshake

WebSocket is not a separate TCP port protocol by default. It reuses the TCP connection and switches protocols with one HTTP/1.1 request:

  1. The client sends GET with Upgrade: websocket, Connection: Upgrade, Sec-WebSocket-Version: 13, and Sec-WebSocket-Key (16 random bytes, Base64).
  2. The server concatenates the key with the magic string 258EAFA5-E914-47DA-95CA-C5AB0DC85B11, computes SHA-1, Base64-encodes the result, and returns it as Sec-WebSocket-Accept with status 101 Switching Protocols.
  3. After 101, the same socket carries WebSocket frames, not HTTP message bodies.

Reverse proxies (Nginx, Cloudflare) must forward Upgrade and Connection end-to-end; otherwise you never get 101 or the connection drops immediately. For restrictive networks, wss:// on port 443 is usually the most reliable option.

Frame structure (summary)

A message may be split across one or more frames. The frame header includes:

  • FIN — whether this frame ends the message (fragmentation possible when FIN is 0).
  • Opcode — e.g. 1 text, 2 binary, 8 close, 9 ping, 10 pong.
  • Payload length — 7-bit length, or extended 16/64-bit fields for larger payloads.
  • Masking — frames from client to server must be XOR-masked per RFC (mitigates cache poisoning). Server to client is not masked.

Application code usually sees reassembled messages, but huge payloads may be fragmented—always set a maximum message size in your stack (read_message_max, server limits, etc.).

Text vs binary frames

Text (opcode 1)Binary (opcode 2)
PayloadUTF-8 textAny bytes (Protobuf, chunks, etc.)
BrowserOften arrives as a string in onmessageBlob / ArrayBuffer
PitfallsInvalid UTF-8 may trigger a protocol error and closeEncoding is your responsibility

JSON chat and notifications usually use text. Protobuf, MsgPack, or file chunks belong in binary. You can mix both on one connection if your app distinguishes frame types.

Heartbeats: Ping/Pong vs app-level keepalive

  • Wire-level — Ping (9) and Pong (10) control frames generate traffic so NATs and load balancers do not treat the TCP session as idle. Pong typically echoes Ping payload.
  • Application-level — JSON {"type":"ping"} or similar is common. Socket.io has its own heartbeat protocol on top of WebSocket.

If the proxy read_timeout is shorter than your heartbeat interval, the link dies regardless of app health. Align Ping period (e.g. 20–30s) with proxy and ALB idle timeouts.

Production patterns

  • WSS on 443 — best pass-through corporate firewalls and proxies.
  • Sticky sessions or shared state — pin a user to an instance when needed; otherwise use Redis Pub/Sub (or streams) to broadcast across nodes.
  • Reconnect — exponential backoff, offline queues, last event / cursor for catch-up after reconnect.
  • Backpressure — do not unboundedly send to slow clients; cap queues and define drop policies.
  • Observability — metrics for open connections, messages/sec, handshake failures, abnormal closes (e.g. 1006 without a close frame).

Native WebSocket API

Server (Node.js)

// server.ts
import { WebSocketServer } from 'ws';
const wss = new WebSocketServer({ port: 8080, maxPayload: 64 * 1024 });
wss.on('connection', (ws) => {
  console.log('Client connected');
  ws.on('message', (data, isBinary) => {
    console.log('Received:', isBinary ? `${data.length} bytes` : data.toString());

    // Echo
    ws.send(`Echo: ${data}`);
  });
  ws.on('close', (code, reason) => {
    console.log('Client disconnected', code, reason.toString());
  });
  ws.send('Welcome!');
});
console.log('WebSocket server running on ws://localhost:8080');

In ws 8.x the message handler receives a Buffer (or an array of buffers) plus an isBinary flag, not a string. Older examples that treat data as a string print <Buffer 48 65 ...> or fail on string methods after an upgrade. maxPayload caps the size of a single message; without a limit, one client can send a very large message and make the server buffer all of it. The default in ws is 100 MiB, which is far more than a chat message needs.

Client (browser)

// client.ts
const ws = new WebSocket('ws://localhost:8080');
ws.onopen = () => {
  console.log('Connected');
  ws.send('Hello Server!');
};
ws.onmessage = (event) => {
  console.log('Received:', event.data);
};
ws.onerror = (error) => {
  console.error('Error:', error);
};
ws.onclose = (event) => {
  console.log('Disconnected', event.code, event.wasClean);
};

The browser API is deliberately small. onerror receives a generic Event with no reason, for security reasons, so the useful information is in onclose: event.code is 1000 for a normal close, 1001 when a page navigates away, and 1006 when the connection died without a close frame (network loss, a proxy timeout, a server crash). 1006 is never sent on the wire; the browser reports it locally. There is also no automatic reconnect: once onclose fires, the object is dead, and a new WebSocket must be created. Calling send() before onopen throws InvalidStateError, which is why the first message above goes inside onopen.

The browser API also cannot set custom headers on the handshake, so Authorization: Bearer ... is not available. The usual options are a cookie (sent automatically for the same site), a short-lived token in the query string (keep it out of access logs), or an authentication message as the first frame, with the server closing the connection if it does not arrive within a few seconds.

Server-side heartbeat with ws

A server cannot tell a quiet client from one whose network vanished: without traffic, the TCP connection can look open for a long time. The ws documentation recommends a ping/terminate loop:

const alive = new WeakMap<WebSocket, boolean>();
wss.on('connection', (ws) => {
  alive.set(ws, true);
  ws.on('pong', () => alive.set(ws, true));
});
const interval = setInterval(() => {
  for (const ws of wss.clients) {
    if (!alive.get(ws)) { ws.terminate(); continue; }
    alive.set(ws, false);
    ws.ping();
  }
}, 30_000);
wss.on('close', () => clearInterval(interval));

Browsers answer ping frames automatically, so nothing is needed on the client. A connection that has not replied by the next round is terminated, which also fires its close handler and lets the application clean up per-user state. Without this, servers slowly accumulate “ghost” connections from phones that switched networks, and connection-count metrics drift upward for no visible reason.


Socket.io

Socket.io is not “WebSocket with helpers”. It is its own protocol (Engine.IO underneath) that usually runs over WebSocket but can fall back to HTTP long-polling, and it adds events with names, acknowledgements, rooms and automatic reconnection. The consequence people hit first: a plain new WebSocket('ws://host:3000') cannot talk to a Socket.io server, and a Socket.io client cannot talk to a plain WebSocket server. Both ends must use Socket.io, with compatible major versions.

Install

# Server
npm install socket.io
# Client
npm install socket.io-client

Server

// server.ts
import express from 'express';
import { createServer } from 'http';
import { Server } from 'socket.io';
const app = express();
const httpServer = createServer(app);
const io = new Server(httpServer, {
  cors: {
    origin: 'http://localhost:5173',   // the page's origin, not the server's
    methods: ['GET', 'POST'],
  },
});
io.on('connection', (socket) => {
  console.log('User connected:', socket.id);
  socket.on('message', (data) => {
    console.log('Message:', data);

    // Pick ONE of these, depending on who should receive it:
    io.emit('message', data);                    // everyone, including the sender
    // socket.broadcast.emit('message', data);   // everyone except the sender
    // io.to(targetSocketId).emit('message', data); // one specific socket
  });
  socket.on('disconnect', (reason) => {
    console.log('User disconnected:', socket.id, reason);
  });
});
httpServer.listen(3000, () => {
  console.log('Server running on :3000');
});

The three emit forms are alternatives. Calling io.emit and socket.broadcast.emit for the same message delivers it twice to everyone except the sender, a bug that is easy to miss when testing with a single browser tab. The cors.origin must be the origin of the page that opens the connection (here a dev server on 5173). With the page and server on the same origin, CORS is not involved at all.

Client

// client.ts
import { io } from 'socket.io-client';
const socket = io('http://localhost:3000');
socket.on('connect', () => {
  console.log('Connected:', socket.id);
});
socket.on('message', (data) => {
  console.log('Received:', data);
});
socket.emit('message', { text: 'Hello!' });

The emit at the end works even before the connection is established, because the Socket.io client buffers events while disconnected and sends them after connecting. That is convenient, but it also means a client that was offline for a minute sends a burst of stale events when it reconnects. For things like “typing” indicators, use socket.volatile.emit(...), which drops the event instead of buffering it.


Example: chat app

Server

// server.ts
import { Server } from 'socket.io';
const io = new Server(3000, {
  cors: { origin: 'http://localhost:5173' },
});
interface Message {
  user: string;
  text: string;
  timestamp: string;
}
io.on('connection', (socket) => {
  console.log('User connected:', socket.id);
  socket.on('join', (room: string) => {
    socket.join(room);
    socket.to(room).emit('user-joined', socket.id);
    console.log(`${socket.id} joined ${room}`);
  });
  socket.on('message', (data: { room: string; message: Message }) => {
    if (!socket.rooms.has(data.room)) return;   // only rooms this socket joined
    io.to(data.room).emit('message', data.message);
  });
  socket.on('typing', (room: string) => {
    socket.to(room).emit('typing', socket.id);
  });
  socket.on('disconnect', () => {
    console.log('User disconnected:', socket.id);
  });
});

The socket.rooms.has check matters more than it looks. Without it, any client can emit { room: 'admins', ... } and post into a room it never joined, because room membership only controls who receives messages, not who may send them. The same caution applies to message.user: it comes from the client and can say anything. In a real application, the server takes the user’s identity from the authenticated handshake (socket.handshake.auth or a session) and overwrites whatever name the client sent.

Client (React)

// Chat.tsx
import { useEffect, useRef, useState } from 'react';
import { io, Socket } from 'socket.io-client';
interface Message {
  user: string;
  text: string;
  timestamp: string;
}
export function Chat() {
  const [messages, setMessages] = useState<Message[]>([]);
  const [input, setInput] = useState('');
  const socketRef = useRef<Socket | null>(null);
  const room = 'general';
  useEffect(() => {
    const socket = io('http://localhost:3000');
    socketRef.current = socket;
    socket.on('connect', () => socket.emit('join', room));  // re-join after every reconnect
    socket.on('message', (message: Message) => {
      setMessages((prev) => [...prev, message]);
    });
    socket.on('typing', (userId: string) => {
      console.log(`${userId} is typing...`);
    });
    return () => {
      socket.disconnect();
    };
  }, []);
  const sendMessage = () => {
    if (input.trim()) {
      const message: Message = {
        user: 'Me',
        text: input,
        timestamp: new Date().toISOString(),
      };
      socketRef.current?.emit('message', { room, message });
      setInput('');
    }
  };
  const handleTyping = () => {
    socketRef.current?.volatile.emit('typing', room);
  };
  return (
    <div>
      <div>
        {messages.map((msg, i) => (
          <div key={i}>
            <strong>{msg.user}:</strong> {msg.text}
          </div>
        ))}
      </div>
      <input
        value={input}
        onChange={(e) => {
          setInput(e.target.value);
          handleTyping();
        }}
        onKeyDown={(e) => e.key === 'Enter' && sendMessage()}
      />
      <button onClick={sendMessage}>Send</button>
    </div>
  );
}

Three details in this component come from real failure modes:

  • Joining inside connect. Rooms live on the server and belong to one socket. When the connection drops and the client reconnects, the server sees a new socket that is in no rooms. If join is emitted once after creating the socket, the chat works until the first network blip and then silently stops receiving messages, while sending still appears to work. Emitting join in the connect handler repeats it after every reconnect.
  • The socket is created and closed inside the effect. In development, React’s Strict Mode mounts, unmounts and mounts again, so a module-level socket created outside the effect ends up connected twice, and every message appears twice. Creating the socket in the effect and disconnecting it in the cleanup keeps exactly one live connection.
  • Typing events are volatile. onChange fires on every keystroke. Sending each one reliably would queue dozens of useless events during a reconnect. In production you would also throttle them to one every second or two.

Messages missed while disconnected are not replayed. Socket.io 4.6 added connection state recovery, which can restore rooms and missed packets after a short disconnection, but it is off by default and does not work with every adapter. For a chat, the robust design is still to store messages and let the client fetch everything after the last message ID it saw when it reconnects.


Redis adapter (scale-out)

import { Server } from 'socket.io';
import { createAdapter } from '@socket.io/redis-adapter';
import { createClient } from 'redis';
const pubClient = createClient({ url: 'redis://localhost:6379' });
const subClient = pubClient.duplicate();
await Promise.all([pubClient.connect(), subClient.connect()]);
const io = new Server(3000, { adapter: createAdapter(pubClient, subClient) });
// Multiple Node processes can now share Socket.io events

With one process, io.to(room).emit() only needs the local list of sockets. With several processes behind a load balancer, a user connected to process A and a user connected to process B are in the “same” room only if the processes share membership information. The Redis adapter publishes every broadcast to Redis, and each process delivers it to its own local sockets. Connecting Redis before creating the server ensures that no broadcast is sent while the adapter is still the in-memory default.

The adapter does not remove the need for sticky sessions. When a client starts with HTTP long-polling (the Socket.io default before upgrading to WebSocket), its requests must reach the same process every time. Without stickiness, requests land on processes that have never heard of the session, and the client loops through 400 Bad Request responses with {"code":1,"message":"Session ID unknown"}. Configure cookie- or IP-based affinity on the load balancer, or start clients with transports: ['websocket'], which skips polling entirely at the cost of losing the fallback for networks that block WebSocket.

This is the scaling bug I would expect first on any Socket.io deployment that moves from one instance to two: everything worked in staging with a single instance, and the “Session ID unknown” errors only appear once a second replica exists, often intermittently, because some requests happen to land on the right process.


FAQ

WebSocket vs Server-Sent Events?

WebSocket is bidirectional. SSE is server → client only, runs over ordinary HTTP, reconnects automatically and resumes with Last-Event-ID. For notifications, dashboards and progress streams, SSE is often simpler to operate; for chat and collaboration, WebSocket (or Socket.io) is the usual choice. See also WebSocket vs SSE vs long polling.

What happens when the connection drops?

Socket.io reconnects with exponential backoff and buffers emits in the meantime. For native WebSocket, implement reconnect with backoff and jitter (so thousands of clients do not reconnect in the same second after a deploy) and state resync (last message ID, snapshot, etc.).

Why does the connection close every 60 seconds behind Nginx?

Nginx’s proxy_read_timeout defaults to 60 seconds, and a WebSocket with no traffic in either direction for that long is closed, which the browser reports as close code 1006. Send pings more often than the proxy timeout, or raise the timeout for the WebSocket location.