Cluster, controller, and metadata
A Kafka cluster is many brokers holding many partition replicas; one broker at a time is the controller that manages metadata, and KRaft (KIP-500) makes that metadata itself a Raft log so the cluster has no external dependency.
ISR membership, leader election, and topic config all need a single source of truth. One broker at a time is the controller — and KRaft made the metadata itself a replicated log, so Kafka has no external dependency.
Scene 04a
Cluster, controller, and metadata
- Watch
- Try it
- Predict
- Capture
Why this matters: a Kafka cluster is many brokers, but somebody has to decide which broker leads which partition, who's in the in-sync replica set, and what topics/configs even exist. That somebody is the controller — a ROLE one broker holds at a time. The top-right badge shows Broker 1 currently holds the controller role. The producer writes a few records to the partition leader (data plane); the controller (control plane) is along for the ride. The big panel at the bottom shows the two ways Kafka has ever stored its metadata side-by-side. KRaft (left card) keeps it INSIDE the Kafka cluster as a Raft log replicated across 3 controllers — one system, one protocol. ZooKeeper (right card) kept it OUTSIDE the cluster in a 3-to-5-node ZK ensemble that you had to operate separately — two systems, two protocols. KIP-500 (2020) replaced ZK with KRaft for exactly the reasons listed on the cards.
Highlighted lines are the ones running in the diagram right now.
def onBrokerFail(brokerId):affected = [p for p in partitionsif p.leader == brokerId]for p in affected:newLeader = electLeader(p)record = PartitionChangeRecord(partition = p.id,leader = newLeader,leaderEpoch = p.epoch + 1,)metadataLog.append(record)# brokers fetch the new record and update their metadata cache
def electLeader(p):for replicaId in p.isr:if replicaId in liveBrokers:return replicaId# ISR is empty — only an out-of-sync replica is leftif unclean.leader.election.enable:return any(p.replicas & liveBrokers)return NO_LEADER # partition goes offline
def append(record): # record is a MetadataRecordif mode == 'kraft':# __cluster_metadata: 3-5 controller votersoffset = raft.appendToQuorum(record)raft.waitForCommit(offset)else: # zookeeperzk.write(path, record.payload)zk.notifyWatchers()# every broker tails the log via the fetch protocolbroadcastToBrokerCaches(record)
Where this sits in Build Kafka
Scene 04a of 13, in the Write side act — Partitioning, replication, and durability knobs.. One controller per cluster; KRaft made the metadata itself a Raft log.
Up next. Durability is four knobs, not one — acks, min.insync.replicas, RF, and unclean.leader.election all interact. Get one wrong and 'acks=all' silently means 'one'.
Designs that use this
- Twitter / X TimelinePush or pull? Both. The canonical fanout problem.
- Uber / Lyft — Match Drivers and RidersMatch a rider to the closest acceptable driver in under 3 s. Geohash, S2, surge.
- Slack / DiscordChannels and history. Push or pull — and how a hot-channel fanout doesn't melt the gateway.
- WhatsApp / MessengerHundreds of millions of long-lived sockets, sub-second 1:1 + group delivery, E2E-encrypted, multi-device, multi-region active-active.