Snapshots — compact without violating consistency

By scene 8 our log keeps growing — every replica has to hold it on disk forever, which is untenable for a long-running cluster. The fix is a snapshot: each replica independently serializes its state machine at lastApplied and discards the log prefix, remembering the boundary in two numbers — lastIncludedIndex and lastIncludedTerm — which stand in for the truncated tail in the AppendEntries consistency check from scene 4. For followers that have fallen below the leader's truncation point, the leader ships the snapshot directly via the InstallSnapshot RPC. None of this violates any of the five invariants — Log Matching keeps working because the snapshot's (lastIncludedIndex, lastIncludedTerm) acts as a synthetic prevLog.

Previously

Scene 7 made cluster membership safe to change: every transitional decision still passes through a majority that overlaps both the old and the new rosters. So now the cluster can grow and shrink without breaking Election Safety. The next operational reality is that the log itself keeps growing — and truncating it has to preserve Log Matching just as carefully as reconfiguration preserved Election Safety.

Scene 08

Snapshots — compact the log without breaking Log Matching

  1. Watch
  2. Try it
  3. Predict
  4. Capture
S1LEADERT:4vote:S1193t4194t4195t4196t4197t4198t4199t4200t4+192S2FOLLOWERT:4vote:—193t4194t4195t4196t4197t4198t4199t4200t4+192S3FOLLOWERT:4vote:—193t4194t4195t4196t4197t4198t4199t4200t4+192term colors: 1 blue · 2 teal · 3 amber · 4 violet · 5 roseAppendEntriesheartbeatRequestVoteInstallSnapshot▲commitIndexNo snapshots yet. Every server has applied 200 commands and the log strip is already overflowing 8 visible cells. Every new commi…
log already overflowing
What to watch for

By scene 8, our log keeps growing. Every replica has to hold it on disk forever — that's untenable for a long-running cluster. The solution: occasionally take a snapshot of the state machine's current state, then throw away the log entries that built up to it. Here are three servers, all caught up at applied index 200, with a log strip that's already overflowing. Watch the next three captions — they install the three new terms this scene needs: snapshot, the (lastIncludedIndex, lastIncludedTerm) pair, and InstallSnapshot RPC.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Replica.takeSnapshot
local operation: serialize state machine, persist metadata, truncate log prefix
def takeSnapshot(self):
# 1. serialize state machine at lastApplied
snap = self.stateMachine.serialize()
snap.lastIncludedIndex = self.lastApplied
snap.lastIncludedTerm = self.log[self.lastApplied].term
# 2. include latest committed configuration in the snapshot
snap.config = self.latestCommittedConfig()
# 3. fsync the snapshot file durably
persist(snap)
# 4. truncate log prefix — entries 1..lastIncludedIndex go away
self.log.discardThrough(snap.lastIncludedIndex)
# NOTE: no RPC, no quorum, no leader involvement.
Leader.replicateTo(follower)
AppendEntries when in-range; InstallSnapshot when below truncation
def replicateTo(self, F):
if self.nextIndex[F] >= self.logStartIndex:
# in-range: ordinary AppendEntries
prev = self.nextIndex[F] - 1
send(AppendEntries(
term=self.currentTerm, prevLogIndex=prev,
prevLogTerm=self.termAt(prev),
entries=self.log[self.nextIndex[F]:],
leaderCommit=self.commitIndex), to=F)
else:
# F has fallen below our truncation point
send(InstallSnapshot(
term=self.currentTerm, leaderId=self.id,
lastIncludedIndex=self.snap.lastIncludedIndex,
lastIncludedTerm=self.snap.lastIncludedTerm,
data=self.snap.bytes, done=True), to=F)
# AppendEntries resume next tick from lastIncludedIndex+1
Follower.handleInstallSnapshot
adopt if past commitIndex; the metadata is the synthetic prevLog
def handleInstallSnapshot(self, msg):
if msg.term < self.currentTerm: return Reject
self.stepDownIfHigherTerm(msg.term)
if msg.lastIncludedIndex <= self.commitIndex:
return Ok # snapshot is older than what we have; ignore
# 1. install the snapshot bytes into the state machine
self.stateMachine.restore(msg.data)
# 2. record the synthetic prevLog: any future AppendEntries with
# prevLogIndex == lastIncludedIndex matches lastIncludedTerm
self.log.resetTo(msg.lastIncludedIndex, msg.lastIncludedTerm)
self.commitIndex = msg.lastIncludedIndex
self.lastApplied = msg.lastIncludedIndex
return Ok

Where this sits in Build Raft — consensus you can defend

Scene 08 of 12. Per-replica snapshots at applied index. (lastIncludedIndex, lastIncludedTerm) substitute for the truncated tail in the AppendEntries consistency check, so Log Matching survives compaction. InstallSnapshot ships the prefix to far-behind followers.

Up next. With safety, reconfiguration, and compaction all defended, the safety machinery is complete. The next question is the read path: with Raft underneath, is it safe for the leader to just serve reads from its local state? It turns out the obvious answer is wrong — naive leader reads are NOT linearizable, even with Raft, because a deposed leader doesn't know it's deposed yet.

All 12 scenes in Build Raft — consensus you can defend · Every curriculum

Built with Arqly
Every scene in Build Raft — consensus you can defend builds on the one before it.All 12 Build Raft — consensus you can defend scenes