Why strong consistency took S3 fourteen years
S3's metadata cache was eventually consistent for 14 years — a write and a read could hit different cache nodes, so a read returned a stale index pointer; the Dec 2020 fix was not to bypass the cache but to make it coherent with a witness that records per-object write order and acts as a read barrier, reloading stale metadata before answering — and it is strong only within one region.
The bytes are durable, but the index that points at them was, for fourteen years, sometimes lying to readers — and making it stop lying was S3's hardest problem.
Scene 08
Why strong consistency took S3 fourteen years
- Watch
- Try it
- Predict
- Capture
Eventual consistency: a PUT updates the index through one cache node, but a GET/LIST served from a DIFFERENT cache node still holds the OLD pointer. Watch three real anomalies replay in turn — each one a reader being told something that isn't true any more.
Highlighted lines are the ones running in the diagram right now.
def read(key):node = cacheRing.nodeFor(key) # may differ from write's nodeentry = node.lookup(key)if not cfg.strong:return entry # may be stale: an older pointer, or a miss# strong: ask the witness whether this entry is currentif witness.isStale(key, entry):entry = persistence.load(key) # reload source of truthnode.put(key, entry) # repair the cache nodereturn entry # read-your-writes, within this region
Where this sits in Build an S3-style distributed object store
Scene 08 of 12, in the Make it real act — Strong consistency, multipart upload, lifecycle & tiering.. Eventual-consistency anomalies, the Dec 2020 flip, and the witness read-barrier — strong within a region only.
Up next. With the index now telling the truth, the last piece of the object's life is handling the big uploads that don't fit in a single request — and they come with a famous integrity gotcha.
All 12 scenes in Build an S3-style distributed object store · Every curriculum