Real-Time Architecture with Socket.io: Lessons from Building Sandesh and Enterprise Portals

September 12, 2024

Real-time features are one of those areas where the gap between "it works in demos" and "it works in production under load" is enormous. I've built real-time systems twice: first as a learning project (Sandesh Lite, a Socket.io messaging app), then for real at Edelta Corporation where live notifications needed to be reliable for enterprise claim workflows.

Here are the patterns and hard lessons from both.

Where I Started: Sandesh Lite

Sandesh was a simple instant messaging app built with React and Socket.io. The architecture was naive but educational:

Architecture Flowchart
Rendering diagram...

Every message fired an event, Firebase stored it, and all connected clients received the broadcast. It worked great for 2–3 concurrent users in development.

In a local demo with 20 concurrent browser tabs, it fell apart. Messages duplicated, rooms got confused, reconnects caused event storms.

That pain taught me to understand Socket.io properly before touching it in production.

The Core Concepts Worth Understanding

Rooms Are Your Namespacing Primitive

// Server side
io.on('connection', (socket) => {
  socket.on('join-claim', (claimId) => {
    socket.join(`claim:${claimId}`);
    console.log(`User ${socket.id} joined claim room ${claimId}`);
  });

  socket.on('claim-status-update', ({ claimId, newStatus, updatedBy }) => {
    // Only broadcast to users watching this specific claim
    io.to(`claim:${claimId}`).emit('status-changed', {
      claimId,
      newStatus,
      updatedBy,
      timestamp: Date.now(),
    });
  });
});

In the enterprise portals, every claims officer connected to only the claim rooms they had open — not all 4,000+ active claims.

Acknowledgements Prevent Lost Events

Don't use fire-and-forget for important events. Socket.io supports acknowledgements:

// Client
socket.emit('update-claim-status', payload, (response) => {
  if (response.error) {
    // Show error UI
    setError(response.error);
  } else {
    // Confirm UI update
    updateLocalClaimState(response.updatedClaim);
  }
});

// Server
socket.on('update-claim-status', async (payload, callback) => {
  try {
    const updatedClaim = await claimService.updateStatus(payload);
    callback({ success: true, updatedClaim });
    // Broadcast to room after confirming write
    io.to(`claim:${payload.claimId}`).emit('status-changed', updatedClaim);
  } catch (err) {
    callback({ error: err.message });
  }
});

This pattern: write to DB first, then broadcast — never the other way around.

Redis Adapter for Horizontal Scaling

A single Socket.io server can handle thousands of connections, but you'll eventually need to scale horizontally. Socket.io's Redis adapter lets multiple server instances share event state:

import { createAdapter } from '@socket.io/redis-adapter';
import { createClient } from 'redis';

const pubClient = createClient({ url: process.env.REDIS_URL });
const subClient = pubClient.duplicate();

await Promise.all([pubClient.connect(), subClient.connect()]);
io.adapter(createAdapter(pubClient, subClient));

Now io.to('claim:123').emit(...) broadcasts to ALL connected clients in that room, regardless of which server instance they're connected to.

Handling Reconnects Gracefully

The nastiest production bug I hit: after a network blip, clients reconnected but missed events that fired during disconnection. Users had stale claim states.

The fix: optimistic UI + re-fetch on reconnect.

// Client
socket.on('connect', () => {
  // Re-fetch current state from REST API
  if (activeClaim) {
    fetchClaimStatus(activeClaim.id).then(updateLocalState);
  }
});

socket.on('disconnect', (reason) => {
  setConnectionStatus('reconnecting');
  if (reason === 'io server disconnect') {
    // Server forced disconnect — re-initiate
    socket.connect();
  }
});

Socket.io's automatic reconnect handles the transport layer. Your job is to handle the application state reconciliation.

React Native: Real-Time Push Notifications

For the field inspector mobile app (React Native), WebSockets worked for in-app real-time. But for background notifications when the app was closed, we used AWS SNS + Firebase Cloud Messaging (FCM):

Architecture Flowchart
Rendering diagram...

In-app: Socket.io room events. Background: Push notifications via SNS/FCM.

The inspector always got notified — regardless of whether they had the app open.

Performance Numbers from Production

ScenarioSetupConcurrent UsersAvg Latency
Sandesh (dev)Single Node.js + Firebase~2050ms
Claims Portal (staging)Single ECS task + no Redis~15080ms
Claims Portal (production)2 ECS tasks + Redis adapter~60035ms

The Redis adapter addition was the most impactful single change — halved latency and removed the hard connection ceiling.

Three Rules for Real-Time in Production

  1. Write to the database first, broadcast second. If the DB write fails after broadcasting, clients have stale state and there's no source of truth.
  2. Re-fetch on reconnect. Don't trust that you received all events during a disconnect. Reconcile with the server REST API.
  3. Use rooms aggressively. Broadcasting to all connected clients is almost always wrong. Broadcast to the room of users who care about that specific data.

Real-time is a force multiplier for user experience, but only if the reliability matches user expectations. Build for reconnects from day one.

GitHub
LinkedIn