The payoff: durability is repair winning a race — scrub, MTTR, and rebuilding lost fragments

Eleven nines is not the fragment count you stamped once — it is the steady state where a background scrub-and-repair loop rebuilds lost fragments faster than disks destroy them, so repair time (MTTR) is the real durability lever.

Previously

Now that the index plane tracks every fragment and the data plane holds them, a background process can watch for losses and fix them — which is exactly the answer to the mystery we left open at the start: how do eleven nines survive disks failing every week?

Scene 07

The payoff: durability is repair winning a race

  1. Watch
  2. Try it
  3. Predict
  4. Capture
DISK FLEET · scrub + repair2,400,000 disksMTTR ~9 hscrub ONDURABILITY11 ninesshowing 180 of 2,400,000 disksdata fragparity fragdeadrottedREDUNDANCY DEFICITbelow full k+m0 / 4 fragments below redundancyOne object, 17 data + 3 parity fragments, scattered across the fleet — watched by a scrub worker and a repair worker.
What to watch for

Disks die on the fleet cadence — every few minutes, somewhere, one fails and takes a fragment with it. Watch the two background workers keep the object whole: the scrub worker sweeps fragments at rest checking they haven't silently gone bad, and the repair worker reads the survivors and rebuilds each lost fragment onto a fresh disk. The deficit gauge keeps drifting back to zero.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Scrub.sweep
the background sweep that catches silent rot at rest
# runs continuously on the background-ops fleet
def sweep():
for frag in fleet.fragments_at_rest():
stored = read(frag.location)
if checksum(stored) != frag.expected_checksum:
# silent bit rot: a live disk, a bad fragment
mark_lost(frag) # hand to Repair
advance_cursor()
Repair.rebuild
reads k survivors, reconstructs the lost fragment
def rebuild(obj, lost_frag):
survivors = read_any(obj.fragments, count = k)
rebuilt = reed_solomon_decode(survivors)
fresh = pick_disk(distinct_failure_domain)
write(fresh, rebuilt) # full k+m restored
index.point(obj, lost_frag, fresh)
# MTTR = wall-clock this rebuild took
Durability.estimate
why repair time, not fragment count, sets the nines
def annual_loss_probability(failure_rate, mttr):
# one fragment down opens a window of length mttr
window = failure_rate * mttr
# need MORE THAN m losses inside one window to lose obj
p_lose = window ** (m + 1)
# m (code) sets the exponent; mttr scales the base
return p_lose

Where this sits in Build an S3-style distributed object store

Scene 07 of 12, in the Keep it alive act — Erasure coding, placement, the two planes, and repair.. Eleven nines is a rate equation: scrub finds rot, repair rebuilds lost fragments faster than disks destroy them.

Up next. The fragments and bytes are now durable — but the index that points at them was, for fourteen years, telling readers the wrong answer, and fixing that was S3's hardest problem.

All 12 scenes in Build an S3-style distributed object store · Every curriculum

Built with Arqly
Every scene in Build an S3-style distributed object store builds on the one before it.All 12 Build an S3-style distributed object store scenes