Real-time features are one of those areas where the gap between "it works in demos" and "it works in production under load" is enormous. I've built real-time systems twice: first as a learning project (Sandesh Lite, a Socket.io messaging app), then for real at Edelta Corporation where live notifications needed to be reliable for enterprise claim workflows.
Here are the patterns and hard lessons from both.
Where I Started: Sandesh Lite
Sandesh was a simple instant messaging app built with React and Socket.io. The architecture was naive but educational:
Every message fired an event, Firebase stored it, and all connected clients received the broadcast. It worked great for 2–3 concurrent users in development.
In a local demo with 20 concurrent browser tabs, it fell apart. Messages duplicated, rooms got confused, reconnects caused event storms.
That pain taught me to understand Socket.io properly before touching it in production.
The Core Concepts Worth Understanding
Rooms Are Your Namespacing Primitive
// Server side
io.on('connection', (socket) => {
socket.on('join-claim', (claimId) => {
socket.join(`claim:${claimId}`);
console.log(`User ${socket.id} joined claim room ${claimId}`);
});
socket.on('claim-status-update', ({ claimId, newStatus, updatedBy }) => {
// Only broadcast to users watching this specific claim
io.to(`claim:${claimId}`).emit('status-changed', {
claimId,
newStatus,
updatedBy,
timestamp: Date.now(),
});
});
});
In the enterprise portals, every claims officer connected to only the claim rooms they had open — not all 4,000+ active claims.
Acknowledgements Prevent Lost Events
Don't use fire-and-forget for important events. Socket.io supports acknowledgements:
// Client
socket.emit('update-claim-status', payload, (response) => {
if (response.error) {
// Show error UI
setError(response.error);
} else {
// Confirm UI update
updateLocalClaimState(response.updatedClaim);
}
});
// Server
socket.on('update-claim-status', async (payload, callback) => {
try {
const updatedClaim = await claimService.updateStatus(payload);
callback({ success: true, updatedClaim });
// Broadcast to room after confirming write
io.to(`claim:${payload.claimId}`).emit('status-changed', updatedClaim);
} catch (err) {
callback({ error: err.message });
}
});
This pattern: write to DB first, then broadcast — never the other way around.
Redis Adapter for Horizontal Scaling
A single Socket.io server can handle thousands of connections, but you'll eventually need to scale horizontally. Socket.io's Redis adapter lets multiple server instances share event state:
import { createAdapter } from '@socket.io/redis-adapter';
import { createClient } from 'redis';
const pubClient = createClient({ url: process.env.REDIS_URL });
const subClient = pubClient.duplicate();
await Promise.all([pubClient.connect(), subClient.connect()]);
io.adapter(createAdapter(pubClient, subClient));
Now io.to('claim:123').emit(...) broadcasts to ALL connected clients in that room, regardless of which server instance they're connected to.
Handling Reconnects Gracefully
The nastiest production bug I hit: after a network blip, clients reconnected but missed events that fired during disconnection. Users had stale claim states.
The fix: optimistic UI + re-fetch on reconnect.
// Client
socket.on('connect', () => {
// Re-fetch current state from REST API
if (activeClaim) {
fetchClaimStatus(activeClaim.id).then(updateLocalState);
}
});
socket.on('disconnect', (reason) => {
setConnectionStatus('reconnecting');
if (reason === 'io server disconnect') {
// Server forced disconnect — re-initiate
socket.connect();
}
});
Socket.io's automatic reconnect handles the transport layer. Your job is to handle the application state reconciliation.
React Native: Real-Time Push Notifications
For the field inspector mobile app (React Native), WebSockets worked for in-app real-time. But for background notifications when the app was closed, we used AWS SNS + Firebase Cloud Messaging (FCM):
In-app: Socket.io room events. Background: Push notifications via SNS/FCM.
The inspector always got notified — regardless of whether they had the app open.
Performance Numbers from Production
| Scenario | Setup | Concurrent Users | Avg Latency |
|---|---|---|---|
| Sandesh (dev) | Single Node.js + Firebase | ~20 | 50ms |
| Claims Portal (staging) | Single ECS task + no Redis | ~150 | 80ms |
| Claims Portal (production) | 2 ECS tasks + Redis adapter | ~600 | 35ms |
The Redis adapter addition was the most impactful single change — halved latency and removed the hard connection ceiling.
Three Rules for Real-Time in Production
- Write to the database first, broadcast second. If the DB write fails after broadcasting, clients have stale state and there's no source of truth.
- Re-fetch on reconnect. Don't trust that you received all events during a disconnect. Reconcile with the server REST API.
- Use rooms aggressively. Broadcasting to all connected clients is almost always wrong. Broadcast to the room of users who care about that specific data.
Real-time is a force multiplier for user experience, but only if the reliability matches user expectations. Build for reconnects from day one.