Live Viewer Count (YouTube/Twitch)

What would you ask before drawing a single box?

Ambiguity you would resolve with the interviewer: scope, scale, who uses it, what counts as done.

About Live Viewer Count (YouTube/Twitch)

Millions of viewers on one entity. Approximate by design.

Difficulty
advanced
Time
about 75 minutes
Stages
10
Topic
Real-Time: Chat, Presence & Live Updates

How this problem is worked

Ten stages, from the questions you would ask an interviewer to the trade-offs you would defend. Each asks one question, and the simulator runs the architecture you draw against the requirements you wrote.

  1. 01ClarificationsWhat would you ask before drawing a single box?
  2. 02Functional reqsWhat must this system actually do?
  3. 03Non-functionalWhat must it promise about speed, uptime and correctness?
  4. 04Capacity estimationHow much load and data does this have to hold?
  5. 05API designWhat does the outside world call, and what comes back?
  6. 06Data modelWhat gets stored, and what is it looked up by?
  7. 07Use-case breakdownHow does each requirement actually get served?
  8. 08High-level designWhich components handle a request, and in what order?
  9. 09Deep divesWhich part breaks first, and what do you do about it?
  10. 10Trade-offsWhat did this design cost, and what breaks at 10×?

Primary sources for this problem

  • Flajolet, Fusy, Gandouet, Meunier — HyperLogLog (DMTCS 2007)
  • Heule, Nunkesser, Hall — HyperLogLog in Practice (EDBT 2013)
  • Cormode, Muthukrishnan — Count-Min Sketch (J.Alg. 2005)
  • Apache DataSketches — Theta sketch (Yahoo)
  • Engineering at Meta — Under the hood: Broadcasting live video to millions (2015)
  • Engineering at Meta — Scaling Live streaming for millions of viewers (2020)
  • Twitch Engineering — State of Engineering 2023 (Spade, PubSub, Kinesis)
  • Twitch Engineering — Breaking the Monolith at Twitch (2022)
  • Twitch Engineering — The QoUX Journey (2025)
  • Twitch Engineering — How Twitch Uses PostgreSQL (2016)
  • Twitch Developers — PubSub API (≤10 conns/IP, ≤50 topics/conn)
  • Discord — How Discord Scaled Elixir to 5,000,000 Concurrent Users (Manifold + Semaphore + FastGlobal)
  • Discord — Real-time Communication at Scale with Elixir (2020)
  • Slack Engineering — Real-time Messaging (Presence Servers + Gateway Servers)
  • Slack Engineering — Migrating Millions of Concurrent WebSockets to Envoy
  • HasGeek Rootconf — Scaling hotstar.com for 25M concurrent viewers (2019)
  • Pragmatic Engineer — Live streaming at world-record scale with Ashutosh Agrawal (JioHotstar)
  • Last9 — Cricket Scale Series #1 (IPL 30M concurrent)
  • ByteByteGo — How Disney+ Hotstar / JioHotstar scales (NAT-per-subnet, multi-CDN)
  • Cloudflare — June 21 2022 cross-region routing retro
  • AWS — Kinesis Data Streams Nov 25 2020 retro (thread-limit cascading failure)
  • Apache Flink — Stateful Stream Processing + Watermarks + RocksDB checkpoints
  • Confluent — KIP-429 Cooperative-Sticky Rebalance
  • Confluent — KIP-794 Strictly Uniform Sticky Partitioner
  • Kreps — Questioning the Lambda Architecture (Kappa, 2014)
  • Beyer et al. — SRE Workbook (Managing Load, Addressing Cascading Failures)
  • YouTube Help — How engagement metrics are counted (the 'we freeze on purpose' rule)

Browse the full problem catalog, or see what the simulator does and does not model.