Disks fail weekly — so what does durable mean? — eleven nines and fleet annual failure rate

Run millions of disks at a 1–4% yearly failure rate and a disk dies every few minutes, so 'durable' cannot mean 'the disk doesn't fail' — it has to mean 'the disk failing doesn't lose your bytes', a designed-for target of eleven nines.

Previously

We cracked the bucket open and asked what 'forever' has to survive; the answer starts at the hardware — and the hardware is millions of disks that are dying continuously.

Scene 02

Disks fail weekly — so what does durable mean?

  1. Watch
  2. Try it
  3. Predict
  4. Capture
DISK FLEET2,400,000 disks · AFR 2%🕑 00:00:00a disk dies every 11 minutesELEVEN NINES TARGETlose ~1 object / 10,000 yrsshowing 200 of 2,400,000 disksalivedeadtracked object (single copy)tracked object: intactMillions of disks. A death every few minutes — a lone copy is doomed. So what does 'durable' mean?
Eleven nines = the designed-for target: ~1 object lost per 10,000 years per 10 million stored. Essentially never.
What to watch for

Watch the fleet. The clock ticks and disks flip from green to dead-red on a steady beat. One outlined tile holds a tracked object — its only copy. Keep watching that tile.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Fleet.deathCadence
the arithmetic the two sliders drive
def death_cadence(fleet_size, afr):
# afr = per-disk annual failure rate (1%..4%)
deaths_per_year = fleet_size * afr
seconds_per_year = 365 * 24 * 3600
# the gap between two disk deaths, in seconds
gap = seconds_per_year / deaths_per_year
return gap
# one disk: gap is months — a death is rare
# millions of disks: gap collapses to minutes
# the per-disk rate never changed; the fleet did
Durability.target
what 'eleven nines' names — a target, not a guarantee
# a designed-for probabilistic target, not a measured SLA
target = 0.99999999999 # eleven nines / year
expected_lost_fraction = 1 - target # 1e-11
# stored 10_000_000 objects?
# expect to lose ~1 object every 10_000 years
objects = 10_000_000
years_per_loss = 1 / (objects * expected_lost_fraction)
# this is DURABLE: do the bytes still EXIST.
# separate from AVAILABLE: can you REACH them now.

Where this sits in Build an S3-style distributed object store

Scene 02 of 12, in the The stakes act — What an object store promises — and the durability problem.. Millions of disks at a few % AFR: one dies every few minutes. Eleven nines — and why copies alone are too costly.

Up next. If surviving disk death means keeping the bytes somewhere other than the dead disk, the very first question is: where do those bytes live, and what keeps track of where they are?

All 12 scenes in Build an S3-style distributed object store · Every curriculum

Built with Arqly
Every scene in Build an S3-style distributed object store builds on the one before it.All 12 Build an S3-style distributed object store scenes