Reads — ReadIndex and lease

Reading from 'the leader' is not automatically correct: a partitioned old leader still inside its election timeout can answer from stale local memory. ReadIndex restores correctness with one heartbeat round (gated on the no-op-on-election); lease reads skip the round in exchange for a clock-skew assumption that GC pauses, fsync stalls, and VM steal can break.

Previously

Scene 8 closed the safety arc — committed entries survive crashes, membership changes are safe, and log truncation preserves Log Matching. The READ path looks like it should already be safe (it's the leader answering, after all) — but the next surprise is that a naive read from 'the leader' silently breaks linearizability under a network partition.

Scene 09

Reads — ReadIndex and lease

  1. Watch
  2. Try it
  3. Predict
  4. Capture
(A) Naive(B) ReadIndex(C) LeaseSTALE READ SERVED — INVARIANT VIOLATEDS1STALE LEADERT:4commit:100applied:100⛓S2LEADERT:5commit:103applied:103✓ no-opS3FOLLOWERT:5commit:103applied:103✓ no-op🚫READ x → S1—clientHere's something that catches even experienced engineers off guard. Reading from the Raft leader sounds safe — but a partitioned …FRESH read · lease green windowSTALE / clock violatedAppendEntries (heartbeat)⛓partitioned server
What to watch for

Here's something that catches even experienced engineers off guard. Suppose you're using etcd to look up a config value, and you ask the leader. The leader answers from its own memory. Sounds safe — it's the leader, right? Now imagine the network partitioned the leader (call it S1) from a majority of the cluster, and a new leader (S2) has already been elected on the other side. Our partitioned 'leader' S1 doesn't know it's been deposed yet — its election timeout hasn't fired, so by its own clock it's still in charge. It happily answers our read with a stale value while a newer write has already been committed by the new leader. We just got a stale-leader read — and we broke linearizability (the property that every read sees the result of every write that completed before it, as if the whole system were a single machine processing one operation at a time). Watch S1 hand the client x=1 while S2 has already committed x=2 on {S2, S3}.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Leader.read (naive)
the broken baseline — read from local state, no quorum check
# NAIVE: leader serves reads from local state with no quorum check.
def on_client_read(key) (leader):
return state_machine.get(key)
# BUG: a partitioned old leader still believes it leads.
# Its election timeout hasn't fired yet, so it has no way to
# know a new leader has already been elected and committed
# a newer write. Stale read served. Linearizability broken.
Leader.read (ReadIndex)
§6.4 commit-barrier read, gated on no-op-on-election
def on_client_read(key) (leader):
# precondition: leader must have committed an entry of its
# CURRENT term (the no-op-on-election). Otherwise commitIndex
# may be stale and ReadIndex would under-report.
if not has_committed_entry_in_current_term():
return defer # buffered in pendingReadIndexMessages
read_index = commit_index # 1. snapshot barrier
acks = { self } # 2. confirm leadership
for peer in cluster_minus_self:
send AppendEntries(heartbeat) -> peer
wait until |acks| >= majority
wait until last_applied >= read_index # 3. apply barrier
return state_machine.get(key) # 4. serve locally
Leader.read (lease)
§6.4.1 trade an RPC round for a clock-bound assumption
# Lease refresh: at every successful heartbeat-round ack.
def on_heartbeat_round_acked() (leader):
lease_expires_at = monotonic_now()
+ election_timeout
- clock_skew_bound
def on_client_read(key) (leader):
if monotonic_now() < lease_expires_at:
return state_machine.get(key) # zero RPCs
else:
return read_index_serve(key) # fall back
# ASSUMPTION: clock skew is bounded. A GC pause, fsync stall,
# or VM steal can let one replica's monotonic clock outrun
# another's — a new leader is elected BEFORE the old leader's
# lease expires from its own perspective, and a stale read
# sneaks through. CockroachDB ties leases to Raft leadership
# specifically to bound this risk.

Where this sits in Build Raft — consensus you can defend

Scene 09 of 12. Naive leader-reads break linearizability under partition. ReadIndex (commit barrier with no-op-on-election precondition) restores it; lease reads buy back the heartbeat round in exchange for a bounded-clock-skew assumption.

Up next. ReadIndex's correctness rests on a tiny one-line mechanism: a no-op log entry the new leader commits at the moment it wins the election. That same one-liner shows up three times across the curriculum and is what makes the new leader's commitIndex trustworthy. Scene 10 puts that mechanism — the no-op-on-election — in the center of the frame.

All 12 scenes in Build Raft — consensus you can defend · Every curriculum

Built with Arqly
Every scene in Build Raft — consensus you can defend builds on the one before it.All 12 Build Raft — consensus you can defend scenes