Leader epoch — the vector clock that fixes truncation

Each record is stamped with a monotonic leader epoch; on restart, followers truncate at the precise epoch boundary instead of a stale High Water Mark, closing the divergence window that pre-KIP-101 HW-based truncation left open.

Previously

ISR explains commit on the happy path. But what happens at failover, when followers must reconcile divergent logs? Pre-KIP-101 used the High Water Mark and silently lost data — the leader epoch is the fix.

Scene 06

Leader epoch — the vector clock that fixes truncation

  1. Watch
  2. Try it
  3. Predict
  4. Capture
Produceracks=allBroker BLEADERLEO=0HW=0LEO=0 · HW=0fetchBroker AHW=0 (lag 1 RTT)IN ISRLEO=0LEO=0
What to watch for

Leader B is at epoch e1. Producer writes m1, m2; follower A fetches them. Watch A's LEO catch up to 2 while its local HW stays at 0 — HW updates piggyback on A's NEXT fetch, so the follower's HW always lags the leader's by one round-trip.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Follower.onBecomeFollower (pre-KIP-101)
the buggy version: truncate to local HW
def onBecomeFollower():
# local HW lags leader HW by one fetch RTT
# — at restart it can be arbitrarily stale
self.log.truncateTo(self.hw)
self.leo = self.hw
# resume fetching from the new leader
loop:
resp = leader.fetch(fromOffset=self.leo)
self.log.append(resp.records)
self.leo += len(resp.records)

Where this sits in Build Kafka

Scene 06 of 13, in the Consensus act — Leader epoch is the vector clock that fixes truncation.. Why HW-based truncation could silently lose acked writes, and how KIP-101 closed the gap.

Up next. Rebalance — when consumers join or leave a group, partitions get reassigned. Eager rebalance stops the world; cooperative-sticky moves only the lanes that actually need to move.

All 13 scenes in Build Kafka · Every curriculum

Built with Arqly
Every scene in Build Kafka builds on the one before it.All 13 Build Kafka scenes