Replication — ISR is not a quorum

A write is committed when every replica currently in the In-Sync Replica set has fetched it — and ISR is mutable, so a slow follower is evicted rather than blocking the high water mark.

Previously

Each partition lives on multiple brokers. The question is: when is a write safe enough to acknowledge? Kafka's answer is the In-Sync Replica set — and it's a moving target, not a fixed quorum vote.

Scene 04

Replication — ISR is not a quorum

  1. Watch
  2. Try it
  3. Predict
  4. Capture
Produceracks=allBroker 1LEADERLEO=0HW=0LEO=0 · HW=0fetchBroker 2caught upIN ISRLEO=0LEO=0fetchBroker 3caught upIN ISRLEO=0LEO=0
What to watch for

Watch the producer append records to the leader. Both followers fetch and their LEOs catch up. The leader's High Water Mark — the offset consumers can read — slides up to the slowest in-ISR follower's LEO, never past it.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Follower.fetch
the fetch loop running on every follower broker
loop forever:
resp = leader.fetch(
partition = p,
fromOffset = self.leo,
)
for record in resp.records:
self.log.append(record)
self.leo += 1
# tells the leader 'I have everything up to leo'
leader.recordFetchPosition(self.id, self.leo)
sleep(replica.fetch.backoff.ms)
Leader.tryAdvanceHW
after every fetch the leader recomputes the high water mark
def tryAdvanceHW():
# ISR is the set of replicas currently caught up.
# 'all' here means all of THESE, not all RF.
inSyncLEOs = [
leo for replicaId, leo in fetchPositions.items()
if replicaId in ISR
]
newHW = min(inSyncLEOs)
if newHW > self.hw:
self.hw = newHW
notifyConsumers(self.hw) # records become readable
Leader.maybeShrinkISR
evicting a slow follower instead of blocking the HW
# runs continuously on the leader
def maybeShrinkISR():
for replicaId in list(ISR):
lastFetch = fetchTimestamps[replicaId]
if now() - lastFetch > replica.lag.time.max.ms:
ISR.remove(replicaId)
controller.notifyISRShrink(
partition, ISR,
)
# HW recomputed against the smaller set
tryAdvanceHW()

Where this sits in Build Kafka

Scene 04 of 13, in the Write side act — Partitioning, replication, and durability knobs.. Why a write commits when the in-sync set fetches it, not a majority.

Up next. Cluster, controller, and metadata — many brokers and many replicas need a coordinator. One broker at a time is the controller, and in KRaft mode the metadata itself is a Raft log.

All 13 scenes in Build Kafka · Every curriculum

Built with Arqly
Every scene in Build Kafka builds on the one before it.All 13 Build Kafka scenes