Pre-Vote and CheckQuorum — fixing the disruptive server

The election scheme from scene 3 has two real liveness bugs; production Raft adds two small mechanisms (one on the candidate, one on the leader) and both are required.

Previously

The election scheme in scene 3 has a real failure mode. Picture a server that briefly loses network access — when it cannot reach the leader it assumes the leader is dead, so it starts an election and bumps its currentTerm. Reconnect it, and its higher term forces the (perfectly healthy) leader to step down (per the universal step-down rule from scene 2). The cluster re-elects for no reason. We call this the disruptive-server bug, and this scene installs the two production fixes for it.

Scene 03a

Pre-Vote and CheckQuorum — fixing the disruptive server

  1. Watch
  2. Try it
  3. Predict
  4. Capture
S1LEADERT:1vote:S11t12t13t14t1S2FOLLOWERT:1vote:—1t12t13t14t1S3FOLLOWERT:1vote:—1t12t13t14t1S4FOLLOWERT:1vote:—1t12t13t14t1S5FOLLOWERT:1vote:—1t12t1PARTITION✓AE ♡ · T1✓AE ♡ · T1✓AE ♡ · T1term colors: 1 blue · 2 teal · 3 amber · 4 violet · 5 roseAppendEntriesheartbeatRequestVoteInstallSnapshot▲commitIndexNo fix turned on. S5 is on the wrong side of the partition. Every time its election timeout fires it bumps its own currentTerm an…
What to watch for

S1 is the leader at term 1, sending heartbeats to S2..S4. S5 is on the wrong side of a network partition — it cannot reach anyone else. Every time S5's election timeout fires (recall from scene 3: the randomized 150–300 ms window after which a follower that has heard nothing from a leader becomes a candidate) it does exactly what scene 3 said to do: it bumps its own currentTerm and tries to start an election. Nobody can answer. Watch its term run away.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Candidate.preVoteProbe / Server.onPreVote
probe peers WITHOUT bumping currentTerm; voter is read-only
# Pre-Vote — added BEFORE the real RequestVote.
on election_timeout:
reset election_timeout
# PROBE: ask peers if they would grant at term+1,
# WITHOUT bumping currentTerm.
pre_votes = { self }
for peer in cluster_minus_self:
send PreVote {
term: currentTerm + 1, # hypothetical
candidateId: self,
lastLogIndex: log.lastIndex,
lastLogTerm: log.lastTerm,
} -> peer
# Only if a majority would grant, bump currentTerm
# and fire the real RequestVote.
on receive PreVote(req):
# No state mutation — this is read-only.
would_grant = req.term > currentTerm
AND isAtLeastAsUpToDate(req.lastLogIndex,
req.lastLogTerm, log)
AND not heard_from_leader_recently()
reply { term: currentTerm, voteGranted: would_grant }

Where this sits in Build Raft — consensus you can defend

Scene 03a of 12. A partitioned member's currentTerm runaway forces a healthy leader to step down on heal. Pre-Vote probes without bumping term; CheckQuorum self-deposes a leader that can't reach a majority. Together they close the §9.6 disruption + the partial-omission liveness hole.

Up next. Now that election liveness is fixed, the leader's actual job begins: it has to copy its log onto every follower. The next scene asks what a single replication message carries — and what tiny consistency check makes it safe to retry forever without ever corrupting a follower's log.

All 12 scenes in Build Raft — consensus you can defend · Every curriculum

Built with Arqly
Every scene in Build Raft — consensus you can defend builds on the one before it.All 12 Build Raft — consensus you can defend scenes