Reads — ReadIndex and lease
Reading from 'the leader' is not automatically correct: a partitioned old leader still inside its election timeout can answer from stale local memory. ReadIndex restores correctness with one heartbeat round (gated on the no-op-on-election); lease reads skip the round in exchange for a clock-skew assumption that GC pauses, fsync stalls, and VM steal can break.
Scene 8 closed the safety arc — committed entries survive crashes, membership changes are safe, and log truncation preserves Log Matching. The READ path looks like it should already be safe (it's the leader answering, after all) — but the next surprise is that a naive read from 'the leader' silently breaks linearizability under a network partition.
Scene 09
Reads — ReadIndex and lease
- Watch
- Try it
- Predict
- Capture
Here's something that catches even experienced engineers off guard. Suppose you're using etcd to look up a config value, and you ask the leader. The leader answers from its own memory. Sounds safe — it's the leader, right?
Now imagine the network partitioned the leader (call it S1) from a majority of the cluster, and a new leader (S2) has already been elected on the other side. Our partitioned 'leader' S1 doesn't know it's been deposed yet — its election timeout hasn't fired, so by its own clock it's still in charge. It happily answers our read with a stale value while a newer write has already been committed by the new leader. We just got a stale-leader read — and we broke linearizability (the property that every read sees the result of every write that completed before it, as if the whole system were a single machine processing one operation at a time). Watch S1 hand the client x=1 while S2 has already committed x=2 on {S2, S3}.
Highlighted lines are the ones running in the diagram right now.
# NAIVE: leader serves reads from local state with no quorum check.def on_client_read(key) (leader):return state_machine.get(key)# BUG: a partitioned old leader still believes it leads.# Its election timeout hasn't fired yet, so it has no way to# know a new leader has already been elected and committed# a newer write. Stale read served. Linearizability broken.
def on_client_read(key) (leader):# precondition: leader must have committed an entry of its# CURRENT term (the no-op-on-election). Otherwise commitIndex# may be stale and ReadIndex would under-report.if not has_committed_entry_in_current_term():return defer # buffered in pendingReadIndexMessagesread_index = commit_index # 1. snapshot barrieracks = { self } # 2. confirm leadershipfor peer in cluster_minus_self:send AppendEntries(heartbeat) -> peerwait until |acks| >= majoritywait until last_applied >= read_index # 3. apply barrierreturn state_machine.get(key) # 4. serve locally
# Lease refresh: at every successful heartbeat-round ack.def on_heartbeat_round_acked() (leader):lease_expires_at = monotonic_now()+ election_timeout- clock_skew_bounddef on_client_read(key) (leader):if monotonic_now() < lease_expires_at:return state_machine.get(key) # zero RPCselse:return read_index_serve(key) # fall back# ASSUMPTION: clock skew is bounded. A GC pause, fsync stall,# or VM steal can let one replica's monotonic clock outrun# another's — a new leader is elected BEFORE the old leader's# lease expires from its own perspective, and a stale read# sneaks through. CockroachDB ties leases to Raft leadership# specifically to bound this risk.
- paperIn Search of an Understandable Consensus Algorithm — §8
- paperConsensus: Bridging Theory and Practice — Chapter 6 (§6.4)
- paperLinearizability: A Correctness Condition for Concurrent Objects
- blogLeader Leases — efficient linearizable reads in CockroachDB
- blogTiKV — Lease Read
- codeetcd-io/raft — ReadIndex + ReadOnly
Where this sits in Build Raft — consensus you can defend
Scene 09 of 12. Naive leader-reads break linearizability under partition. ReadIndex (commit barrier with no-op-on-election precondition) restores it; lease reads buy back the heartbeat round in exchange for a bounded-clock-skew assumption.
Up next. ReadIndex's correctness rests on a tiny one-line mechanism: a no-op log entry the new leader commits at the moment it wins the election. That same one-liner shows up three times across the curriculum and is what makes the new leader's commitIndex trustworthy. Scene 10 puts that mechanism — the no-op-on-election — in the center of the frame.
All 12 scenes in Build Raft — consensus you can defend · Every curriculum