Snapshots — compact without violating consistency
By scene 8 our log keeps growing — every replica has to hold it on disk forever, which is untenable for a long-running cluster. The fix is a snapshot: each replica independently serializes its state machine at lastApplied and discards the log prefix, remembering the boundary in two numbers — lastIncludedIndex and lastIncludedTerm — which stand in for the truncated tail in the AppendEntries consistency check from scene 4. For followers that have fallen below the leader's truncation point, the leader ships the snapshot directly via the InstallSnapshot RPC. None of this violates any of the five invariants — Log Matching keeps working because the snapshot's (lastIncludedIndex, lastIncludedTerm) acts as a synthetic prevLog.
Scene 7 made cluster membership safe to change: every transitional decision still passes through a majority that overlaps both the old and the new rosters. So now the cluster can grow and shrink without breaking Election Safety. The next operational reality is that the log itself keeps growing — and truncating it has to preserve Log Matching just as carefully as reconfiguration preserved Election Safety.
Scene 08
Snapshots — compact the log without breaking Log Matching
- Watch
- Try it
- Predict
- Capture
By scene 8, our log keeps growing. Every replica has to hold it on disk forever — that's untenable for a long-running cluster. The solution: occasionally take a snapshot of the state machine's current state, then throw away the log entries that built up to it. Here are three servers, all caught up at applied index 200, with a log strip that's already overflowing. Watch the next three captions — they install the three new terms this scene needs: snapshot, the (lastIncludedIndex, lastIncludedTerm) pair, and InstallSnapshot RPC.
Highlighted lines are the ones running in the diagram right now.
def takeSnapshot(self):# 1. serialize state machine at lastAppliedsnap = self.stateMachine.serialize()snap.lastIncludedIndex = self.lastAppliedsnap.lastIncludedTerm = self.log[self.lastApplied].term# 2. include latest committed configuration in the snapshotsnap.config = self.latestCommittedConfig()# 3. fsync the snapshot file durablypersist(snap)# 4. truncate log prefix — entries 1..lastIncludedIndex go awayself.log.discardThrough(snap.lastIncludedIndex)# NOTE: no RPC, no quorum, no leader involvement.
def replicateTo(self, F):if self.nextIndex[F] >= self.logStartIndex:# in-range: ordinary AppendEntriesprev = self.nextIndex[F] - 1send(AppendEntries(term=self.currentTerm, prevLogIndex=prev,prevLogTerm=self.termAt(prev),entries=self.log[self.nextIndex[F]:],leaderCommit=self.commitIndex), to=F)else:# F has fallen below our truncation pointsend(InstallSnapshot(term=self.currentTerm, leaderId=self.id,lastIncludedIndex=self.snap.lastIncludedIndex,lastIncludedTerm=self.snap.lastIncludedTerm,data=self.snap.bytes, done=True), to=F)# AppendEntries resume next tick from lastIncludedIndex+1
def handleInstallSnapshot(self, msg):if msg.term < self.currentTerm: return Rejectself.stepDownIfHigherTerm(msg.term)if msg.lastIncludedIndex <= self.commitIndex:return Ok # snapshot is older than what we have; ignore# 1. install the snapshot bytes into the state machineself.stateMachine.restore(msg.data)# 2. record the synthetic prevLog: any future AppendEntries with# prevLogIndex == lastIncludedIndex matches lastIncludedTermself.log.resetTo(msg.lastIncludedIndex, msg.lastIncludedTerm)self.commitIndex = msg.lastIncludedIndexself.lastApplied = msg.lastIncludedIndexreturn Ok
Where this sits in Build Raft — consensus you can defend
Scene 08 of 12. Per-replica snapshots at applied index. (lastIncludedIndex, lastIncludedTerm) substitute for the truncated tail in the AppendEntries consistency check, so Log Matching survives compaction. InstallSnapshot ships the prefix to far-behind followers.
Up next. With safety, reconfiguration, and compaction all defended, the safety machinery is complete. The next question is the read path: with Raft underneath, is it safe for the leader to just serve reads from its local state? It turns out the obvious answer is wrong — naive leader reads are NOT linearizable, even with Raft, because a deposed leader doesn't know it's deposed yet.
All 12 scenes in Build Raft — consensus you can defend · Every curriculum