Socket.IO v4 in Production: Rooms, Acks, Missed Events and Redis Scale-Out

Key takeaways

Socket.IO is a protocol on top of WebSocket and HTTP long-polling, not WebSocket itself. This guide covers the v4 API (rooms, acks with timeouts, volatile emits, auth middleware) and the production problems that bite later: CORS, missed events during reconnects, sticky sessions and the Redis adapter.

What Socket.IO actually is

Socket.IO gives you event-based, bidirectional messaging between a browser (or any client) and a Node.js server. The part that trips people up early is that Socket.IO is not WebSocket. It is a protocol layered on two transports:

  • Engine.IO, the low-level layer, which opens an HTTP long-polling connection first (by default) and then tries to upgrade to WebSocket. It handles heartbeats (ping/pong) and detects dead connections.
  • Socket.IO, the layer on top, which adds named events, namespaces, rooms, acknowledgements, and automatic reconnection on the client.

That layering is where both the convenience and the costs come from. You get reconnection, rooms and request/response semantics without writing them yourself. In exchange, both sides must speak the Socket.IO protocol, every message carries a little extra framing, and horizontal scaling needs more infrastructure than a plain WebSocket server would.

If you’re still choosing between WebSocket, Server-Sent Events and polling, read WebSocket vs SSE vs Long Polling first. For the underlying RFC 6455 protocol (handshake, frames, ping/pong), see Real-Time Apps with WebSocket and Socket.io. This article is about Socket.IO v4 specifically.

Proof: a raw WebSocket client can’t talk to it

I ran this against a Socket.IO 4.8 server using the ws package as the client:

import WebSocket from 'ws';

// 1. The obvious attempt: the server's root URL
new WebSocket('ws://localhost:3100/')
  .on('error', (e) => console.log(e.message)); // "socket hang up"

// 2. The Engine.IO endpoint connects, but the payload isn't your events
const w = new WebSocket('ws://localhost:3100/socket.io/?EIO=4&transport=websocket');
w.on('message', (m) => console.log(m.toString()));
// 0{"sid":"n-RNw2loaHJpiO5EAAAA","upgrades":[],"pingInterval":25000,...}

The first attempt gets hung up on because nothing listens for upgrades at /. The second one connects, but it gets an Engine.IO “open” packet. To send an event you’d have to build 42["message",{...}] frames by hand and answer pings, or the server drops you after the ping timeout. So if your consumers are IoT devices, a mobile team that wants a standard WebSocket, or a third-party integration, a plain ws server is usually the better fit. Socket.IO makes sense when you control both ends and want what it adds on top.

Setup and a minimal server

npm install socket.io express     # server
npm install socket.io-client      # client
import express from 'express';
import { createServer } from 'node:http';
import { Server } from 'socket.io';

const app = express();
const httpServer = createServer(app);

const io = new Server(httpServer, {
  cors: {
    origin: ['http://localhost:5173', 'https://app.example.com'],
    credentials: true, // only if you rely on cookies
  },
});

io.on('connection', (socket) => {
  console.log('connected', socket.id);

  socket.on('message', (data) => {
    socket.emit('reply', { text: 'Got it', at: Date.now() });
  });

  socket.on('disconnect', (reason) => {
    console.log('disconnected', socket.id, reason);
  });
});

httpServer.listen(3000);

Attach Socket.IO to the http.Server, not to the Express app. app.listen() creates its own server internally, and if you call it instead of httpServer.listen(), the Socket.IO handler is attached to a server nobody is listening on. The client then gets 404s on /socket.io/.

CORS since v3

Since v3, Socket.IO has CORS off by default. A frontend on another origin (including another port on localhost) needs an explicit cors option. Without it, the browser console shows something like:

Access to XMLHttpRequest at 'http://localhost:3000/socket.io/?EIO=4&transport=polling&t=...'
from origin 'http://localhost:5173' has been blocked by CORS policy

On the client, all you see is a connect_error with the generic message xhr poll error. Two details matter:

  • origin: '*' cannot be combined with credentials: true. Browsers reject a wildcard origin on credentialed requests, so list your origins explicitly.
  • The client also needs withCredentials: true if you want cookies sent on the polling requests.

cors only affects the HTTP long-polling requests. Browsers don’t enforce CORS on the WebSocket upgrade. If you need to restrict which sites can open WebSocket connections, check socket.handshake.headers.origin in middleware or use the allowRequest option.

The client and its reconnection model

import { io } from 'socket.io-client';

const socket = io('http://localhost:3000', {
  auth: (cb) => cb({ token: localStorage.getItem('token') }),
});

socket.on('connect', () => console.log('connected', socket.id));

socket.on('connect_error', (err) => {
  // err.message is whatever the server middleware passed to next(new Error(...))
  console.log('connect_error', err.message, 'will retry:', socket.active);
});

socket.on('disconnect', (reason) => {
  if (reason === 'io server disconnect') {
    // The server kicked us via socket.disconnect(): no automatic retry
    socket.connect();
  }
  // 'transport close', 'ping timeout' etc. -> the client reconnects on its own
});

Three behaviors are worth knowing before production:

  1. The client reconnects automatically, with exponential backoff, except when the server calls socket.disconnect() (reason io server disconnect) or when middleware rejects the connection. In those cases socket.active is false and you decide whether to retry. I checked this directly: after a middleware rejection, connect_error fired with active= false.
  2. Client-side emits are buffered while disconnected and flushed on reconnect. That’s convenient for chat, but a burst of stale events can arrive all at once. Use socket.volatile.emit() for data that’s worthless when late, or check socket.connected before emitting.
  3. Using auth as a function matters for token refresh. The function runs on every reconnection attempt, so a token rotated in localStorage gets picked up. A static object would keep sending the expired token forever.

socket.id is regenerated on every reconnection, so never use it as a user identifier. Put the user ID in socket.data from your auth middleware and join a room named after it (see below).

Rooms and broadcasting

Rooms are server-side groupings of sockets. The client doesn’t know which rooms it’s in; it only receives what the server sends.

io.on('connection', (socket) => {
  const userId = socket.data.user.id;
  socket.join(`user:${userId}`); // every tab/device of this user

  socket.on('join-room', (roomId) => {
    // validate that this user may join roomId before trusting it
    socket.join(roomId);
    socket.to(roomId).emit('user-joined', { userId });
  });

  socket.on('room-message', ({ roomId, text }) => {
    if (!socket.rooms.has(roomId)) return; // don't let clients post into rooms they're not in
    io.to(roomId).emit('new-message', { from: userId, text, at: Date.now() });
  });
});

// Anywhere else in the server (e.g. after an HTTP request):
io.to(`user:${someUserId}`).emit('notification', { text: 'Order shipped' });
CallWho receives it
socket.emit()This socket only
socket.broadcast.emit()Everyone in the namespace except this socket
io.emit()Everyone in the namespace
io.to(room).emit()Everyone in the room
socket.to(room).emit()Everyone in the room except this socket
io.except(room).emit()Everyone except members of the room

Sockets are removed from all rooms automatically on disconnect, and they are not rejoined on reconnect unless you use connection state recovery. That is why the “per-user room joined in the connection handler” pattern is sturdier than joining rooms in response to client events: it re-runs on every connection. For client-requested rooms, have the client re-emit join-room in its connect handler.

A security note: join-room with a client-supplied ID is an authorization decision. If you don’t check it, any client can subscribe to any room by guessing its name.

Namespaces

Namespaces split one server into independent channels. Each has its own connection handler, its own rooms and its own middleware. On the client they share one underlying connection (multiplexing).

const admin = io.of('/admin');

admin.use(authenticate); // the same JWT middleware as the main namespace (see below)
admin.use((socket, next) => {
  if (socket.data.user.role !== 'admin') return next(new Error('forbidden'));
  next();
});

admin.on('connection', (socket) => {
  socket.on('kick-user', (userId) => {
    io.in(`user:${userId}`).disconnectSockets();
  });
});
const socket = io('https://api.example.com');            // main namespace "/"
const adminSocket = io('https://api.example.com/admin'); // same connection, "/admin"

Middleware registered with io.use() applies to the main namespace only. If you add an /admin namespace, it needs its own admin.use(...), including auth. Otherwise it’s open to anyone who connects to it. In practice rooms cover most “separate channel” needs. I’d use a namespace when the auth rules or the event vocabulary genuinely differ.

Acknowledgements, timeouts, and volatile emits

Plain emits give you no delivery guarantee. Acknowledgements add request/response semantics: the receiver calls a callback, and the sender gets the result.

// Server
socket.on('save-message', async (message, callback) => {
  try {
    const saved = await db.messages.insert(message);
    callback({ ok: true, id: saved.id });
  } catch (err) {
    callback({ ok: false, error: 'save failed' });
  }
});

// Client: promise style with a timeout (v4.6+)
try {
  const res = await socket.timeout(5000).emitWithAck('save-message', { text: 'Hello' });
  console.log(res.ok ? `saved ${res.id}` : res.error);
} catch (err) {
  // err.message === "operation has timed out"
  showRetryButton();
}

Always pair acks with timeout(). Without it, if the server handler throws before calling the callback or the connection drops, emitWithAck returns a promise that never settles. With timeout(), the callback style also receives an error as its first argument: socket.timeout(5000).emit('x', data, (err, res) => ...). I verified both forms; the rejection message is operation has timed out.

An ack means your handler ran and called back. That’s still at-most-once delivery: if the ack is lost and the client retries, the server may process the message twice. Put a client-generated ID on messages that must not be duplicated, and dedupe on the server.

At the other extreme, volatile emits are for data where freshness matters more than completeness:

socket.volatile.emit('cursor', { x, y }); // dropped if the socket isn't ready to write

Live cursors, typing indicators and game positions are good candidates. If you replay a hundred stale cursor positions after a reconnect, the UI visibly jumps.

Some names are reserved: connect, connect_error, disconnect, disconnecting, newListener, and removeListener. Calling socket.emit('disconnect') throws "disconnect" is a reserved event name.

Authentication in middleware

import jwt from 'jsonwebtoken';

function authenticate(socket, next) {
  const token = socket.handshake.auth?.token;
  if (!token) return next(new Error('unauthorized'));

  try {
    socket.data.user = jwt.verify(token, process.env.JWT_SECRET);
    next();
  } catch {
    const err = new Error('invalid token');
    err.data = { code: 'TOKEN_EXPIRED' }; // arrives on the client as err.data
    next(err);
  }
}

io.use(authenticate);

Middleware runs once per connection, not per event. That’s efficient, but it means a long-lived socket keeps working after the token expires, until the socket reconnects. If revocation matters (logout everywhere, banned users), disconnect the user’s sockets when that happens: io.in(user:${id}).disconnectSockets(true). Or check an expiry timestamp stored in socket.data inside sensitive event handlers. Prefer handshake.auth over query strings for tokens, because query strings end up in proxy access logs.

Missed events and connection state recovery

This is the gap that surprises people most. When a client drops off the network for a few seconds, every server-to-client event emitted during that window is gone. The client reconnects, gets a new socket.id, is in no rooms, and silently misses whatever happened in the meantime.

Socket.IO 4.6 added connection state recovery:

const io = new Server(httpServer, {
  connectionStateRecovery: {
    maxDisconnectionDuration: 2 * 60 * 1000, // how long the server keeps state
    skipMiddlewares: true,                    // skip auth on a successful recovery
  },
});

io.on('connection', (socket) => {
  if (socket.recovered) {
    // id, rooms and socket.data were restored and missed packets were replayed
  } else {
    // fresh session: join rooms, send an initial snapshot
  }
});

I tested this by closing the client’s underlying transport, emitting to its room while it was offline, and letting it reconnect. With recovery enabled, the client got the missed event and socket.recovered was true on both sides. The volatile emit sent during the same window was not replayed, which is exactly right.

The limits are the important part:

  • Server restarts lose the state, because it’s kept in the adapter (memory, by default). A deploy disconnects everyone, and nobody gets recovered.
  • Adapter support varies. The built-in in-memory adapter supports it, and so do some others (the Redis Streams adapter and the MongoDB adapter, for example). The classic Redis pub/sub adapter does not. Check the adapter’s docs before relying on it.
  • It is best-effort. For a chat history or an order feed, the reliable design is still to persist messages with a monotonic ID and have the client fetch “everything after ID X” on reconnect. Recovery then just smooths over brief blips.

The failure I’ve seen most with Socket.IO is exactly this one. Everything works on a stable office network. Then mobile users on flaky connections report that “messages sometimes don’t show up until I refresh.” Nothing is broken server-side and no error is logged, because a dropped emit to a disconnected socket isn’t an error. Once you realize that Socket.IO’s reconnection restores the connection but not the data, the fix is usually a catch-up fetch on the client’s connect handler, and then the reports stop.

Scaling horizontally: sticky sessions and an adapter

Running more than one Node.js process requires two separate things. Teams often set up one and forget the other.

Sticky sessions (because of long-polling)

With the default transports, the handshake starts over HTTP long-polling: several HTTP requests that share one session ID. If a load balancer spreads them round-robin, a request reaches a process that never issued that ID, and the client gets HTTP 400 with:

{"code":1,"message":"Session ID unknown"}

The client keeps reconnecting in a loop. Fix it with session affinity at the load balancer. With nginx:

upstream socketio_nodes {
    ip_hash;                    # same client IP -> same backend
    server 127.0.0.1:3001;
    server 127.0.0.1:3002;
}

server {
    listen 80;
    location /socket.io/ {
        proxy_pass http://socketio_nodes;
        proxy_http_version 1.1;
        proxy_set_header Upgrade $http_upgrade;
        proxy_set_header Connection "upgrade";
        proxy_set_header Host $host;
        proxy_read_timeout 60s;  # must exceed pingInterval + pingTimeout
    }
}

ip_hash breaks down when many users share an IP (corporate NAT), so cookie-based affinity is better where your load balancer supports it. The alternative is to skip long-polling entirely with io(url, { transports: ['websocket'] }). Then each session is one TCP connection and needs no stickiness, but you lose the fallback for networks that block WebSocket. For the upgrade headers and timeouts in general, see Nginx Reverse Proxy Configuration.

An adapter (because of broadcasts)

Sticky sessions keep one client on one process, but io.to('room').emit() on process A only reaches sockets connected to process A. An adapter forwards broadcasts between processes. The Redis adapter uses Redis pub/sub:

npm install @socket.io/redis-adapter redis
import { createClient } from 'redis';
import { createAdapter } from '@socket.io/redis-adapter';

const pubClient = createClient({ url: 'redis://localhost:6379' });
const subClient = pubClient.duplicate();
await Promise.all([pubClient.connect(), subClient.connect()]);

io.adapter(createAdapter(pubClient, subClient));

With ioredis instead, create clients with new Redis(...) (ioredis exports a Redis class, not createClient) and pass them the same way. Once the adapter is in place, io.to(room).emit(), fetchSockets() and disconnectSockets() work across all instances.

Keep in mind that Redis pub/sub is fire-and-forget. If an instance loses its Redis connection briefly, broadcasts sent during that time never reach its clients. That’s one more reason to have the persisted catch-up path. For how Redis pub/sub behaves more broadly, see the Redis guide.

The multi-instance setup gives a second classic failure: it works perfectly on one dev machine, then after scaling to two containers, “half the users don’t get notifications.” Everything passes locally because one process never needs an adapter. When I’ve traced this kind of bug, the cause was almost always that either the adapter was never configured or the ingress wasn’t sticky. The Session ID unknown line in the browser’s network tab points to the second.

React: one socket, stable listeners

The common useRef hook pattern has two bugs: the component never re-renders when the ref is set, and socket.off('event') without a handler removes every listener for that event, including ones registered by other components. A simpler approach is one module-level socket:

// socket.js
import { io } from 'socket.io-client';
export const socket = io(import.meta.env.VITE_API_URL, { autoConnect: false });
// ChatRoom.jsx
import { useEffect, useState } from 'react';
import { socket } from './socket';

export function ChatRoom({ roomId }) {
  const [messages, setMessages] = useState([]);

  useEffect(() => {
    const onMessage = (msg) => setMessages((prev) => [...prev, msg]);
    const onConnect = () => socket.emit('join-room', roomId); // rejoin after reconnects

    socket.on('new-message', onMessage);
    socket.on('connect', onConnect);
    if (socket.connected) onConnect();
    else socket.connect();

    return () => {
      socket.off('new-message', onMessage); // remove only this listener
      socket.off('connect', onConnect);
      socket.emit('leave-room', roomId);
    };
  }, [roomId]);

  const send = (text) => socket.emit('room-message', { roomId, text });

  return (
    <ul>
      {messages.map((m, i) => <li key={i}>{m.text}</li>)}
    </ul>
  );
}

In development, React StrictMode mounts, unmounts and remounts effects. If your effect creates a socket and the cleanup doesn’t close it, you’ll see two connections per tab and every message rendered twice. Keeping the socket outside the component avoids that, and so does always passing the handler to off().

When not to use Socket.IO

  • One-way server-to-client streams (notifications, live logs, LLM token streaming): Server-Sent Events work over plain HTTP, reconnect automatically with Last-Event-ID, and need no special client library.
  • Clients you don’t control or non-JavaScript consumers expecting standard WebSocket: use a plain WebSocket server.
  • Guaranteed, ordered delivery between services: that’s a message broker’s job (Kafka, RabbitMQ). Socket.IO is for talking to clients, not for backend-to-backend messaging.

Where Socket.IO pays off is interactive apps where you control both ends: chat, collaborative editing presence, dashboards with user-to-user interaction, multiplayer lobbies. Rooms, acks and reconnection handling there would otherwise be several hundred lines of your own code.