Leader election — vote, majority, restriction
To become leader of a term, a server raises its hand by bumping the term, asks every other server 'will you vote for me?', and wins only if MORE THAN HALF say yes — and a voter only says yes if it hasn't already voted in this term AND the asker's log isn't behind its own.
Scene 2 gave us a term — a monotonically-increasing integer that segments time and makes any server reset to follower the moment it sees a number bigger than its own. The obvious next question: which server is the leader of the current term? Each term needs at most one, and the cluster has to pick one without a coordinator. That picking process is leader election.
Scene 03
Leader election — vote, majority, restriction
- Watch
- Try it
- Predict
- Capture
Now that we have a term that monotonically segments time, the obvious question is: which server is the leader of the current term? A follower that hears nothing from a leader for a while gives up waiting and tries to become leader itself. The 'while' is a random wait — somewhere between 150 and 300 ms in the standard configuration — and we call it the election timeout: the randomized interval after which a follower that has received no contact from a leader transitions to candidate and begins an election. Watch: S1's timeout fires first. It bumps its currentTerm from 1 to 2, marks itself with votedFor=S1 (the persisted record of who, if anyone, this server has voted for in the current term — at most one per term), and sends every peer a 'will you vote for me?' message. The systems term for that message is a RequestVote RPC — short for Remote Procedure Call, which just means 'a function call sent over the network'; the call carries the candidate's term, its id, and the position of its last log entry so the voter can decide whether to grant. Four yes-replies come back; combined with its own self-vote that is 5 out of 5, comfortably above the majority quorum (more than half of the cluster — for 5 servers, that's 3). S1 becomes leader of term 2 and immediately sends an empty append (a heartbeat) to tell every follower it is in charge.
Highlighted lines are the ones running in the diagram right now.
on election_timeout: # randomized 150–300 msrole = candidatecurrentTerm += 1votedFor = selfpersist(currentTerm, votedFor)votes = { self }reset election_timeoutfor peer in cluster_minus_self:send RequestVote {term: currentTerm,candidateId: self,lastLogIndex: log.lastIndex,lastLogTerm: log.lastTerm,} -> peer
on receive RequestVote(req):# (universal step-down already handled at handler top — see scene 2)up_to_date = isAtLeastAsUpToDate(req.lastLogIndex, req.lastLogTerm, log,)# PLACEHOLDER — precise predicate is sharpened in scene 6 (§5.4.1).if (votedFor == null OR votedFor == req.candidateId) AND up_to_date:votedFor = req.candidateIdpersist(currentTerm, votedFor)reply { term: currentTerm, voteGranted: true }else:reply { term: currentTerm, voteGranted: false }
on receive RequestVoteReply(reply, from peer):if reply.term > currentTerm: # (step-down at handler top)returnif reply.voteGranted:votes.add(peer)if |votes| >= majority(cluster.size): # ⌊N/2⌋ + 1role = leaderinitialize nextIndex[p] = log.lastIndex + 1 for p in clusterinitialize matchIndex[p] = 0 for p in clusterbroadcast empty AppendEntries (heartbeat) # claim leadership
Where this sits in Build Raft — consensus you can defend
Scene 03 of 12. Election timeout → candidate → RequestVote → majority grants → leader. Voters grant only if the candidate's log is at-least-as-up-to-date — a placeholder predicate refined precisely in scene 6.
Up next. We just bought election safety (at most one leader per term) — but at a real cost: any server that briefly loses contact, increments its term a few times, then reconnects can force the current healthy leader to step down even though that server can never win the vote. That is the disruptive-server bug, and the next scene plugs it.
All 12 scenes in Build Raft — consensus you can defend · Every curriculum