Build your own Kafka
A partitioned, replicated, append-only log. The log is the database — internalize that, and a dozen product designs get easier.
- Scenes
- 13 interactive scenes
- Time
- about 91 minutes
- Topic
- Queues, Pub/Sub & Event Streaming
What you are building, and why
You are designing a distributed log: producers append messages, consumers read them in order, the system survives machine failures. This is the canonical "log is the database" system — once you have built it, you understand why partitions matter for ordering, why exactly-once is hard, what ISR means when a broker dies, and why log compaction has the trade-offs it does.
Resist the urge to "describe Kafka." Make decisions yourself, defend them, and let the AI push back. The point isn't to recreate Kafka byte-for-byte — it's to make every choice that the Kafka authors made, and feel why they made it.
What you will be able to explain afterwards
- append-only log
- partitioning
- replication
- ISR & quorum
- log compaction
- exactly-once semantics
- controller / KRaft
Why a log?
Orientation — the log is the database, not a queue.
- 01Foundations — what Kafka is, words you'll hearWhy Kafka exists and the seven core nouns (producer, broker, topic, partition, record, offset, consumer). Orientation before you touch anything.~7 min
- 01aHello Kafka — topic, brokers, recordsFoundations: what's a topic vs a partition, what's a broker, what does the producer/consumer code actually look like.~7 min
- 02The log is the database — per-consumer offsets on an append-only logWhy a log isn't a queue, and why that one fact unlocks the rest.~7 min
- 02aOffsets, retention, and where bookmarks liveRead and commit are separate ack channels; retention, not consumers, ages records out.~7 min
Write side
Partitioning, replication, and durability knobs.
- 03Partitions — splitting the logParallelism by sharding ordering. Hot partitions, key skew.~7 min
- 04Replication — ISR is not a quorumWhy a write commits when the in-sync set fetches it, not a majority.~7 min
- 04aCluster, controller, and metadataOne controller per cluster; KRaft made the metadata itself a Raft log.~7 min
- 05Durability is four knobs, not one — acks=all, min.insync.replicas and unclean.leader.electionacks, min.insync.replicas, RF, unclean — and how 'all' silently means one.~7 min
- 05aLog compaction — keep the last value per keyCompact-retention turns the log into a state store; tombstones propagate deletes.~7 min
Consensus
Leader epoch is the vector clock that fixes truncation.
Scale
Rebalancing without halting every consumer.
Guarantees
Three monotonic counters; one design canvas.
Where you'll use this
Product designs whose trade-offs turn on what this curriculum teaches.
- Twitter / X TimelinePush or pull? Both. The canonical fanout problem.
- Uber / Lyft — Match Drivers and RidersMatch a rider to the closest acceptable driver in under 3 s. Geohash, S2, surge.
- Slack / DiscordChannels and history. Push or pull — and how a hot-channel fanout doesn't melt the gateway.
- WhatsApp / MessengerHundreds of millions of long-lived sockets, sub-second 1:1 + group delivery, E2E-encrypted, multi-device, multi-region active-active.
- Instagram News FeedRanked feed with cursor pagination. No `OFFSET`.
- Collaborative Editor (Google Docs)OT vs CRDT. Causal ordering. Real conflict-freedom. Server-authoritative single-writer doc-actor with WAL-before-ack — per-doc serialization, persisted before broadcast, pinned to a home region.
More in Queues, Pub/Sub & Event Streaming
Moving events between services without losing them: logs, work queues, fanout, change data capture, stream processing, and the delivery systems built on top.
- Build Build a Message Queue (RabbitMQ / SQS)A point-to-point work queue — the messaging primitive Kafka is NOT. Each message goes to one consumer, ack deletes, retries push to a dead-letter queue, and a poisoned message is everyone's problem. Internalize ack vs visibility timeout vs DLQ vs prefetch vs FIFO groups — and learn to tell when Kafka is the wrong tool and when a queue is.
- Build Build a CDC pipeline (Debezium + outbox)Your service writes to its DB and publishes to Kafka — and any crash between those two writes is permanent inconsistency. Build a Change Data Capture pipeline (modeled on Debezium + the outbox pattern) that closes the gap by making the database itself the event source.
- Ad Click AggregatorStream processing with watermarks and exactly-once.
- Notification SystemPush, email, SMS. Idempotent. Failover.
- Distributed Cron — Mass Scheduled EmailSingle trigger, 50M recipient idempotency, catch-up.
Prefer to design it yourself?
The same subject as a staged workspace: draw the architecture, and a simulator traces requests through the boxes you drew.
Open the Build Kafka workspace