Durability is four knobs, not one — acks=all, min.insync.replicas and unclean.leader.election

acks=all only guarantees no loss when min.insync.replicas >= 2 and unclean.leader.election=false; otherwise ISR can shrink to {leader} and 'all' silently means one.

Previously

ISR decides commit, the controller decides leadership. Now the four knobs that decide what 'durable' actually means — and how a wrong combination turns 'acks=all' into 'acks=one' without any error.

Scene 05

Durability is four knobs, not one

  1. Watch
  2. Try it
  3. Predict
  4. Capture
Producersend(record)Brokeracks=all · MIR=2 · RF=3 · unclean=offPARTITIONS · 3LeaderLEADER · in ISRFollower 1follower · in ISRFollower 2follower · in ISRDOWNSTREAM CONSUMERwhat would survive a leader crash right nowConsumerreads up to HW (safe)
What to watch for

Default config: RF=3, min.insync.replicas=2, acks=all, unclean.leader.election=off. Producer sends, all three replicas append, leader acks. The downstream reader sees a fully replicated log — a leader crash here loses zero acked writes.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Producer.send(record, acks=all)
the producer-side wait loop — block until the broker acks
def send(record):
leader = metadata.leaderFor(record.partition)
resp = leader.produce(record, acks=cfg.acks)
if cfg.acks == 0:
return # fire-and-forget, no wait
if cfg.acks == 1:
return resp # leader-only ack
# acks == 'all': leader replies only after the
# broker-side commit gate has accepted the record.
return resp
Broker.onProduce(record)
the commit gate — reject before silently degrading
def onProduce(record):
self.log.append(record) # leader appends first
if cfg.acks != 'all':
return Ack() # no ISR check on acks=0/1
if len(ISR) < cfg.min_insync_replicas:
raise NotEnoughReplicas(
isr=len(ISR), required=cfg.min_insync_replicas,
)
waitForFetch(record.offset, replicas=ISR)
self.hw = min(leo[r] for r in ISR)
return Ack()
Controller.maybeUncleanElect(partition)
the gate that lets a stale replica win an empty-ISR election
def maybeUncleanElect(partition):
if len(ISR) > 0:
return promote(pickFromISR())
if not cfg.unclean_leader_election_enable:
partition.state = UNAVAILABLE
return None # KIP-106 default
candidate = pickAnyLiveReplica()
candidate.epoch += 1 # bumps leader epoch
return promote(candidate)

Where this sits in Build Kafka

Scene 05 of 13, in the Write side act — Partitioning, replication, and durability knobs.. acks, min.insync.replicas, RF, unclean — and how 'all' silently means one.

Up next. Log compaction — keep the last value per key. Compacted topics turn the log itself into a state store, and underpin __consumer_offsets, Streams k-tables, and CDC sinks.

All 13 scenes in Build Kafka · Every curriculum

Built with Arqly
Every scene in Build Kafka builds on the one before it.All 13 Build Kafka scenes