Hold the write until it wakes up — hinted handoff and the hint TTL

Hinted handoff lets a coordinator absorb a short replica outage by stashing missed writes locally and replaying them when the replica returns — but only within a TTL.

Previously

Partitions are dramatic; the everyday case is a single replica blipping for a few seconds — for that the cluster has a softer trick.

Scene 10

Hold the write until it wakes up

  1. Watch
  2. Try it
  3. Predict
  4. Capture
mode: with-hintABCDE (do…DEADuser-42R1R2R3HINT BUFFERcoordinator stash for EHINTS DISK0% fullAGE / TTL0.0h / 3.0h→ replay on replica returnbest-effort — coordinator-only.coordinator stashes a hint per missed write; replays on E's return.
What to watch for

Replica E just went DOWN. Watch the simulated outage clock tick — for every missed write, the coordinator stashes a hint envelope. The age meter creeps toward the 3h TTL. Within TTL, the envelope replays when E returns; past it, the envelope dissolves.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Coordinator.put
the write path with the dead-replica hint branch
def put(key, value):
replicas = ring.replicas_for(key)
for r in replicas:
if r.alive:
send_async(r, Write(key, value))
else:
# stash on local disk for later replay
hint_store.append(
Hint(target=r, write=(key, value),
created_at=now()),
)
return Ack # to W replicas (best-effort hints)
HintStore.replay_on_return
fires when the down replica gossips back as alive
def replay_on_return(replica):
for hint in hint_store.for(replica):
age = now() - hint.created_at
if age <= max_hint_window_in_ms:
send(replica, hint.write)
hint_store.remove(hint)
else:
# past TTL — replica returns 'stale forever'
hint_store.remove(hint)
HintStore.sweep
background TTL sweep — discard hints past max_hint_window
# runs continuously on the coordinator
def sweep():
for hint in list(hint_store):
age = now() - hint.created_at
if age > max_hint_window_in_ms:
hint_store.remove(hint) # silent loss
sleep(hints_flush_period_ms)

Where this sits in Build a wide-column store (Cassandra / DynamoDB family)

Scene 10 of 13, in the Healing act — Hinted handoff + read repair + anti-entropy + gossip keep the cluster honest.. When a replica is briefly unreachable, the coordinator stashes the write and replays it on return.

Up next. Hints expire and replicas drift cold; the cluster needs a way to find and fix divergence on its own.

All 13 scenes in Build a wide-column store (Cassandra / DynamoDB family) · Every curriculum

Built with Arqly
Every scene in Build a wide-column store (Cassandra / DynamoDB family) builds on the one before it.All 13 Build a wide-column store (Cassandra / DynamoDB family) scenes