Cluster, controller, and metadata

A Kafka cluster is many brokers holding many partition replicas; one broker at a time is the controller that manages metadata, and KRaft (KIP-500) makes that metadata itself a Raft log so the cluster has no external dependency.

Previously

ISR membership, leader election, and topic config all need a single source of truth. One broker at a time is the controller — and KRaft made the metadata itself a replicated log, so Kafka has no external dependency.

Scene 04a

Cluster, controller, and metadata

  1. Watch
  2. Try it
  3. Predict
  4. Capture
Producertopic: eventsBroker 1LEADERLEO=0HW=0LEO=0 · HW=0fetchBroker 2followerISR ?LEO=0LEO=0fetchBroker 3followerISR ?LEO=0LEO=0ControllerKRaftbroker-1Failover candidatesBroker 2Broker 3METADATA LAYER — KRaft vs ZooKeeper (this is what KIP-500 replaced)currently: KRaftKRaft (KIP-500, modern)INSIDE the Kafka clusterKafka clusterC1C2C3Raft replication (same process tree)1 system to operateKafka only — no separate ensembleSub-second failovercontroller hop in ~millisecondsScales to millionsof partitions per clusterSingle protocolRaft consensus inside KafkaZooKeeper (legacy, pre-KIP-500)OUTSIDE the Kafka clusterZK ensembleKafkawatchertwo clusters, two ops teams2 systems to operateKafka + a 3–5 node ZK ensembleSeconds-level failovermetadata reload from ZK~200k partition ceilingZK watcher storms at high countsTwo protocolsKafka wire + ZAB (ZK's protocol)
What to watch for

Why this matters: a Kafka cluster is many brokers, but somebody has to decide which broker leads which partition, who's in the in-sync replica set, and what topics/configs even exist. That somebody is the controller — a ROLE one broker holds at a time. The top-right badge shows Broker 1 currently holds the controller role. The producer writes a few records to the partition leader (data plane); the controller (control plane) is along for the ride. The big panel at the bottom shows the two ways Kafka has ever stored its metadata side-by-side. KRaft (left card) keeps it INSIDE the Kafka cluster as a Raft log replicated across 3 controllers — one system, one protocol. ZooKeeper (right card) kept it OUTSIDE the cluster in a 3-to-5-node ZK ensemble that you had to operate separately — two systems, two protocols. KIP-500 (2020) replaced ZK with KRaft for exactly the reasons listed on the cards.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Controller.onBrokerFail
the active controller reacts when a broker stops heartbeating
def onBrokerFail(brokerId):
affected = [
p for p in partitions
if p.leader == brokerId
]
for p in affected:
newLeader = electLeader(p)
record = PartitionChangeRecord(
partition = p.id,
leader = newLeader,
leaderEpoch = p.epoch + 1,
)
metadataLog.append(record)
# brokers fetch the new record and update their metadata cache
Controller.electLeader
pick a new leader from replicas still in the ISR
def electLeader(p):
for replicaId in p.isr:
if replicaId in liveBrokers:
return replicaId
# ISR is empty — only an out-of-sync replica is left
if unclean.leader.election.enable:
return any(p.replicas & liveBrokers)
return NO_LEADER # partition goes offline
MetadataLog.append
KRaft: metadata changes are a Raft log on the controller quorum
def append(record): # record is a MetadataRecord
if mode == 'kraft':
# __cluster_metadata: 3-5 controller voters
offset = raft.appendToQuorum(record)
raft.waitForCommit(offset)
else: # zookeeper
zk.write(path, record.payload)
zk.notifyWatchers()
# every broker tails the log via the fetch protocol
broadcastToBrokerCaches(record)

Where this sits in Build Kafka

Scene 04a of 13, in the Write side act — Partitioning, replication, and durability knobs.. One controller per cluster; KRaft made the metadata itself a Raft log.

Up next. Durability is four knobs, not one — acks, min.insync.replicas, RF, and unclean.leader.election all interact. Get one wrong and 'acks=all' silently means 'one'.

All 13 scenes in Build Kafka · Every curriculum

Built with Arqly
Every scene in Build Kafka builds on the one before it.All 13 Build Kafka scenes