Backend cheatsheetSocket.IO serverOne answer per request
Socket.IO
Server
Every emit target, and one rule for answers: each request event replies with event_name:response or event_name:error. Typed, validated and the same in plain Socket.IO and in NestJS.
A typed server
Attach Socket.IO to an HTTP server, lock CORS to your app and type every event, including its response and error twin.
A client first opens an HTTP long polling session, then upgrades to WebSocket when it can. Socket.IO then sends its own CONNECT packet with your auth data, runs your middleware and only then fires connection.
Picture it: from first request to connection
export interface Requests {
game_join: { requestId: string; gameId: string };
move_submit: { requestId: string; gameId: string; word: string };
}
export interface Results {
game_join: { players: number };
move_submit: { score: number };
}
export type Ok<T> = { ok: true; event: string; requestId: string; data: T };
export type Fail = { ok: false; event: string; requestId: string; error: { code: string; message: string } };
// every request event gets a :response and an :error twin, checked by the compiler
export type ServerEvents =
{ [K in keyof Requests as `${K & string}:response`]: (res: Ok<Results[K]>) => void } &
{ [K in keyof Requests as `${K & string}:error`]: (err: Fail) => void } &
{ game_state: (state: GameState) => void };
export type ClientEvents = {
[K in keyof Requests]: (req: Requests[K], ack?: (res: Ok<Results[K]> | Fail) => void) => void;
};
Why it matters: the naming rule is enforced by the type system, so a missing error event or a typo is a compile error.
const http = createServer();
export const io = new Server<ClientEvents, ServerEvents, {}, { userId: string }>(http, {
cors: { origin: ['https://app.example.com'], credentials: true },
pingInterval: 25_000, // the defaults, shown so you know they exist
pingTimeout: 20_000,
maxHttpBufferSize: 1e5, // 100 KB is plenty for game moves
});
http.listen(3000);
Why it matters: a small buffer limit stops one client from pushing megabytes into a handler.
Emit targets
Where an event goes depends only on what you call emit on. Here is every common target and exactly who receives it.
Picture it: the cluster every row below refers to
A is the socket whose handler is running. The adapter carries emits from node 1 to node 2.
| Call from A’s handler | A | B | C | D | E | F |
|---|---|---|---|---|---|---|
socket.emit('x') | Yes | No | No | No | No | No |
io.emit('x') | Yes | Yes | Yes | Yes | Yes | No |
socket.broadcast.emit('x') | No | Yes | Yes | Yes | Yes | No |
io.to('game:42').emit('x') | Yes | Yes | No | Yes | No | No |
socket.to('game:42').emit('x') | No | Yes | No | Yes | No | No |
io.except('game:42').emit('x') | No | No | Yes | No | Yes | No |
io.to(socketD.id).emit('x') | No | No | No | Yes | No | No |
io.of('/admin').emit('x') | No | No | No | No | No | Yes |
io.local.emit('x') | Yes | Yes | Yes | No | No | No |
The room, without the sender. The usual way to relay a move.
See the exampleEveryone in the room, sender included.
See the exampleEvery tab and device of one user, through a room you join on connect.
See the exampleMay be dropped when the client is busy. Fine for cursors.
See the exampleio.on('connection', (socket) => {
socket.join(`user:${socket.data.userId}`);
socket.on('move_submit', async (req, ack) => {
const state = await games.apply(req);
socket.to(`game:${req.gameId}`).emit('game_state', state); // the others
io.to(`user:${state.nextPlayer}`).emit('game_state', state); // all their tabs
});
});
Why it matters: user rooms survive reconnects, while socket ids change every time.
Response and error
Every request event answers with event_name:response or event_name:error, in one envelope shape, carrying the caller’s request id.
Pick one rule and use it everywhere. When a client emits move_submit, the server always answers that socket with exactly one of move_submit:response or move_submit:error. Both carry ok, the event name and the requestId the client sent, so a client can match answers to requests even when several are in flight. If the client passed an ack callback, it receives the same envelope.
Picture it: one request, one answer
export function on<K extends keyof Requests>(
socket: AppSocket, event: K, schema: ZodType<Requests[K]>,
run: (req: Requests[K]) => Promise<Results[K]>,
) {
socket.on(event, async (payload: unknown, ack?: (r: Ok<Results[K]> | Fail) => void) => {
const requestId = (payload as { requestId?: string })?.requestId ?? randomUUID();
const reply = typeof ack === 'function' ? ack : () => {};
const fail = (code: string, message: string) => {
const err: Fail = { ok: false, event, requestId, error: { code, message } };
socket.emit(`${event}:error`, err);
reply(err);
};
const parsed = schema.safeParse(payload);
if (!parsed.success) return fail('INVALID_PAYLOAD', parsed.error.issues[0].message);
try {
const res: Ok<Results[K]> = { ok: true, event, requestId, data: await run(parsed.data) };
socket.emit(`${event}:response`, res);
reply(res);
} catch (e) {
e instanceof DomainError ? fail(e.code, e.message) : fail('INTERNAL', 'Could not complete the request');
}
});
}
Why it matters: validation errors, domain errors and crashes all leave in the same shape, and internal messages never reach clients.
The same rule in plain JavaScript, with a fake socket that prints what it emits. Run it to see all three outcomes.
const socket = { handlers: {}, on(e, fn) { this.handlers[e] = fn; },
emit(e, body) { console.log(e.padEnd(21), JSON.stringify(body)); } };
function on(event, validate, run) {
socket.on(event, async (req = {}) => {
const base = { event, requestId: req.requestId };
const problem = validate(req);
if (problem) return socket.emit(event + ':error', { ok: false, ...base, error: problem });
try {
socket.emit(event + ':response', { ok: true, ...base, data: await run(req) });
} catch (e) {
socket.emit(event + ':error', { ok: false, ...base, error: { code: e.code || 'INTERNAL' } });
}
});
}
const words = new Set(['mango', 'sushi']);
on('move_submit',
(req) => (typeof req.word === 'string' ? null : { code: 'INVALID_PAYLOAD' }),
async (req) => {
if (!words.has(req.word)) throw Object.assign(new Error(), { code: 'WORD_UNKNOWN' });
return { score: req.word.length * 2 };
});
(async () => {
await socket.handlers.move_submit({ requestId: 'r1', word: 'sushi' });
await socket.handlers.move_submit({ requestId: 'r2', word: 'pizza' });
await socket.handlers.move_submit({ requestId: 'r3', word: 42 });
})();
$ node contract-model.js move_submit:response {"ok":true,"event":"move_submit","requestId":"r1","data":{"score":10}} move_submit:error {"ok":false,"event":"move_submit","requestId":"r2","error":{"code":"WORD_UNKNOWN"}} move_submit:error {"ok":false,"event":"move_submit","requestId":"r3","error":{"code":"INVALID_PAYLOAD"}}
Why it matters: three different failures and one success, and the client only ever has to handle two event names per request.
| Rule | Why |
|---|---|
Request events are snake_case nouns and verbs | Easy to read in logs, no clash with the : suffix |
Exactly one of :response or :error per request | Clients can await one of two events and time out otherwise |
requestId is echoed back | Parallel requests of the same kind stay matched |
error.code is stable, message is for people | Clients branch on codes, never on text |
Broadcasts to others use their own event, like game_state | A response is for the caller only |
Acks and timeouts
An ack turns an event into a call with a reply. Give every one a timeout so nobody waits forever.
Reply to the caller. Send the same envelope as the response event.
See the exampleAsk a client and await its answer. Needs v4.6 or later.
See the exampleFails the ack with an error if no answer comes in time.
See the exampleAsk a whole room; the callback gets every reply.
See the example// ask one client to confirm, with a deadline
try {
const res = await socket.timeout(5_000).emitWithAck('rematch_offer', { gameId });
if (res.ok) await games.rematch(gameId);
} catch {
socket.emit('rematch_offer:error', { ok: false, event: 'rematch_offer', requestId, error: { code: 'TIMEOUT', message: 'No answer in 5s' } });
}
// ask a room, collect every reply
io.to(`game:${gameId}`).timeout(3_000).emit('ready_check', {}, (err, replies) => {
if (err) log.warn('some players did not answer');
const ready = replies.filter((r) => r.ok).length;
});
Why it matters: the server never hangs on a client that closed the tab, and the timeout itself follows the error contract.
| Use | When |
|---|---|
| Ack | Only the caller needs the result, and needs it to carry on |
event:response | Clients that prefer listeners over callbacks, or reconnect logic that replays |
| A broadcast event | Other people also need to know, like the new game state |
Rooms and namespaces
Rooms are server side groups you join and leave freely. Namespaces are separate channels a client connects to on purpose.
Adds the socket to rooms. Only the server decides, never the client.
See the exampleA Set with the socket id and every joined room.
See the exampleSockets in a room across every node, with their data.
See the exampleKick a whole room or user from anywhere.
See the exampleon(socket, 'game_join', JoinSchema, async ({ gameId }) => {
await games.assertPlayer(gameId, socket.data.userId); // throws DomainError('NOT_A_PLAYER')
await socket.join(`game:${gameId}`);
const players = (await io.in(`game:${gameId}`).fetchSockets()).length;
return { players };
});
socket.on('disconnecting', () => {
for (const room of socket.rooms) if (room.startsWith('game:')) io.to(room).emit('player_left', { userId: socket.data.userId });
});
// one namespace per tenant, created on first connect
io.of(/^\/tenant-\w+$/).on('connection', (s) => { const tenant = s.nsp.name.slice(8); });
Why it matters: joining runs through the same contract, so a refused join answers with game_join:error instead of silence.
Auth middleware
Check the token once, during the handshake, and refuse the connection before any handler runs.
- Client connectsio(url, { auth: { token } })
- io.use()verify the token
- next()socket.data.userId set
- connectionhandlers attach
io.use(async (socket, next) => {
try {
const { sub, exp } = await verifyJwt(socket.handshake.auth.token);
socket.data.userId = sub;
setTimeout(() => socket.disconnect(true), exp * 1000 - Date.now()); // tokens expire, sockets must too
next();
} catch {
const err = new Error('unauthorized') as Error & { data?: unknown };
err.data = { code: 'TOKEN_INVALID' }; // same code vocabulary as event:error
next(err); // the client gets connect_error with message and data
}
});
Why it matters: unauthenticated sockets never reach a handler, and the client can tell an expired token from a network failure.
| Disconnect reason | What happened | Client reconnects? |
|---|---|---|
| transport close | Tab closed or network dropped | Yes |
| ping timeout | No pong in pingInterval plus pingTimeout | Yes |
| server namespace disconnect | You called socket.disconnect() | No |
| client namespace disconnect | The client disconnected itself | No |
| transport error | For example a message over maxHttpBufferSize | Yes |
| server shutting down | The server is closing | Yes |
NestJS gateways
The same contract in a Nest gateway: an interceptor emits event:response, a filter emits event:error.
@Injectable()
export class WsContractInterceptor implements NestInterceptor {
intercept(ctx: ExecutionContext, next: CallHandler) {
const ws = ctx.switchToWs();
const event = ws.getPattern();
const requestId = ws.getData()?.requestId ?? randomUUID();
return next.handle().pipe(map((data) => {
const res = { ok: true, event, requestId, data };
ws.getClient<Socket>().emit(`${event}:response`, res);
return res; // also becomes the ack
}));
}
}
Why it matters: handlers just return data; the envelope and event name are added in one place.
@Catch()
export class WsContractFilter implements ExceptionFilter {
catch(err: unknown, host: ArgumentsHost) {
const ws = host.switchToWs();
const event = ws.getPattern();
const code = err instanceof WsException ? String(err.getError()) : 'INTERNAL';
ws.getClient<Socket>().emit(`${event}:error`, {
ok: false, event, requestId: ws.getData()?.requestId, error: { code, message: 'Request failed' },
});
}
}
Why it matters: Nest’s default filter emits a generic exception event; this keeps the event_name:error rule.
@UseFilters(WsContractFilter)
@UseInterceptors(WsContractInterceptor)
@UsePipes(new ValidationPipe({ whitelist: true, exceptionFactory: () => new WsException('INVALID_PAYLOAD') }))
@WebSocketGateway({ namespace: '/game', cors: { origin: process.env.WEB_ORIGIN } })
export class GameGateway {
@WebSocketServer() server: Namespace;
@SubscribeMessage('move_submit')
async submit(@ConnectedSocket() s: Socket, @MessageBody() dto: MoveDto) {
const state = await this.games.apply(s.data.userId, dto);
this.server.to(`game:${dto.gameId}`).except(s.id).emit('game_state', state);
return { score: state.lastScore };
}
}
Why it matters: with a namespace set, the injected server is a Namespace, so every emit stays inside /game.
Scale past one node
A second pod breaks two things at once: broadcasts only reach local clients, and polling requests land on the wrong node.
Picture it: two nodes, one room
export class RedisIoAdapter extends IoAdapter {
private adapter: ReturnType<typeof createAdapter>;
async connect() {
const pub = createClient({ url: process.env.REDIS_URL });
const sub = pub.duplicate();
await Promise.all([pub.connect(), sub.connect()]);
this.adapter = createAdapter(pub, sub);
}
createIOServer(port: number, options?: ServerOptions) {
const server = super.createIOServer(port, options);
server.adapter(this.adapter);
return server;
}
}
Why it matters: every emit is published to Redis, so a player on node 2 sees a move made on node 1.
upstream socket_nodes {
ip_hash; # same client, same node
server app1:"n">3000;
server app2:"n">3000;
}
location /socket.io/ {
proxy_pass http://socket_nodes;
proxy_http_version "n">1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_read_timeout "n">75s; # longer than pingInterval
}
Why it matters: polling sends several HTTP requests per session and they must all reach the node that holds it.
| Setting | Default | Set it when |
|---|---|---|
pingInterval | 25000 ms | A proxy cuts idle connections sooner |
maxHttpBufferSize | 1 MB | Always: size it to your largest message |
connectionStateRecovery | off | Short drops should restore rooms and missed events. Not supported by the classic Redis adapter; use the Redis Streams adapter |
transports: ['websocket'] | polling first | Client side, to skip polling and sticky sessions |
Which one do I need?
Match what you want to happen to the call that does it, and keep every answer in the response or error shape.
| I want to | Use | Answer with |
|---|---|---|
| Answer the caller | socket.emit(`${e}:response`) plus ack | { ok: true, event, requestId, data } |
| Report a failure to the caller | socket.emit(`${e}:error`) plus ack | { ok: false, event, requestId, error } |
| Tell the rest of a room | socket.to(room).emit() | A named broadcast event |
| Reach every tab of a user | io.to(`user:${id}`) | A named broadcast event |
| Wait for a client’s answer | socket.timeout(ms).emitWithAck() | The same envelope back |
| Refuse a connection | next(err) in io.use | err.data.code |
| Run on more than one pod | Redis adapter plus sticky sessions | Nothing changes for clients |