12
Slack / Discord
Channels and history. Push or pull — and how a hot-channel fanout doesn't melt the gateway.SavedSaved on this device — Saved on this device
01Clarifications
What would you ask before drawing a single box?
Ambiguity you would resolve with the interviewer: scope, scale, who uses it, what counts as done.
AI staff engineer
Enter to send · Shift+Enter for a new line
About Slack / Discord
Channels and history. Push or pull — and how a hot-channel fanout doesn't melt the gateway.
- Difficulty
- intermediate
- Time
- about 50 minutes
- Stages
- 10
- Topic
- Real-Time: Chat, Presence & Live Updates
How this problem is worked
Ten stages, from the questions you would ask an interviewer to the trade-offs you would defend. Each asks one question, and the simulator runs the architecture you draw against the requirements you wrote.
- 01ClarificationsWhat would you ask before drawing a single box?
- 02Functional reqsWhat must this system actually do?
- 03Non-functionalWhat must it promise about speed, uptime and correctness?
- 04Capacity estimationHow much load and data does this have to hold?
- 05API designWhat does the outside world call, and what comes back?
- 06Data modelWhat gets stored, and what is it looked up by?
- 07Use-case breakdownHow does each requirement actually get served?
- 08High-level designWhich components handle a request, and in what order?
- 09Deep divesWhich part breaks first, and what do you do about it?
- 10Trade-offsWhat did this design cost, and what breaks at 10×?
Primary sources for this problem
- Discord — How Discord Stores Trillions of Messages (Cassandra → ScyllaDB)
- Discord — How Discord Scaled Elixir to 5,000,000 Concurrent Users
- Discord — Maxjourney: 1M+ Online in a Single Server (relay tier)
- Discord — Using Rust to Scale Elixir for 11M Concurrent Users (SortedSet NIF)
- Slack — Flannel: an Application-Level Edge Cache to Make Slack Scale
- Slack — Real-time messaging (Channel Server + Gatewayserver)
- Slack — Scaling Datastores at Slack with Vitess
- Slack — Slack's Outage on January 4th, 2021 (TGW saturation)
- Slack — Slack's Incident on 2-22-22 (Consul / Vitess feedback)
- Slack — A Terrible, Horrible, No-Good, Very Bad Day (May 2020 HAProxy)
- Slack — Migration to a Cellular Architecture (InfoQ 2024)
- Slack — Tracing Notifications & How Slack Rebuilt Notifications
- Slack — Migrating Millions of Concurrent WebSockets to Envoy
- elixir-lang.org — Real-Time Communication at Scale with Elixir at Discord (2020)
- Cloudflare — July 17 2020 BGP outage post-mortem
- Google SRE Workbook, Ch. 5 (alerting on SLOs)
Build the primitives this design leans on
Each one is an animated curriculum that constructs the system from scratch.
- Build Build KafkaA partitioned, replicated, append-only log. The log is the database — internalize that, and a dozen product designs get easier.
- Build Build a distributed search engine (Elasticsearch / OpenSearch style)Five million books, a search box, and a 100 ms budget. Build the engine from the inverted index up — segment, refresh, shard, replica, scatter-gather, BM25 — and feel why every guarantee that lives across shards is paid for in either an extra round trip or a small lie about the rankings.
- Build Build RedisAn in-memory data-structure server: one thread, rich types, optional persistence, async replication. Internalize the cost of single-threaded simplicity and a dozen caching/HA decisions get easier.
More in Real-Time: Chat, Presence & Live Updates
Long-lived connections and the things you push down them — messages, cursors, green dots, scores, prices and bids.
- WhatsApp / MessengerHundreds of millions of long-lived sockets, sub-second 1:1 + group delivery, E2E-encrypted, multi-device, multi-region active-active.
- Live Comments / Score UpdatesPub/sub at scale. WebSocket vs SSE vs long polling. Approximate by design — mega-rooms drop comments on purpose.
- Collaborative Editor (Google Docs)OT vs CRDT. Causal ordering. Real conflict-freedom. Server-authoritative single-writer doc-actor with WAL-before-ack — per-doc serialization, persisted before broadcast, pinned to a home region.
- Online IndicatorGreen dot for contacts. Mind the N² watch problem. Approximate by design — never quote presence more precisely than reality.
- Concurrent Hotel Viewers"X users viewing this right now." Hot keys, HLL.
- Live Viewer Count (YouTube/Twitch)Millions of viewers on one entity. Approximate by design.
Browse the full problem catalog, or see what the simulator does and does not model.