Build your own S3-style distributed object store
Eleven nines of durability over disks that fail weekly. Build the object store from first principles — flat keyspace, immutable objects, erasure coding instead of replication, eventual consistency turned strong, multipart upload, lifecycle and tiering — and feel why every modern data lake sits on top of something shaped exactly like this.
- Scenes
- 12 interactive scenes
- Time
- about 84 minutes
- Topic
- Object Storage, File Sync & Media Delivery
What you will be able to explain afterwards
- flat keyspace + bucket + object semantics (no real directories)
- immutable objects + versioning + delete markers
- replication (3x) vs erasure coding (k+m, Reed-Solomon)
- durability math: P(loss) under independent disk failures
- read-after-write consistency (the 2020 strong-consistency flip)
- multipart upload + per-part ETag + completion manifest
- range reads + partial-object cache at the edge
- lifecycle policies: hot → infrequent → archive → delete
- storage classes: Standard / IA / Glacier (cost vs retrieval latency)
- cross-region replication + the strongly-vs-eventually consistent trade
- pre-signed URLs + IAM-style request authorization
- the per-prefix throughput limit (avoid hot prefixes via sharded keys)
The stakes
What an object store promises — and the durability problem.
- 01Hello, object store: it's just a bucketBucket, object, key, PUT, GET — the flat key→bytes model, before we earn the word "forever."~7 min
- 02Disks fail weekly — so what does durable mean? — eleven nines and fleet annual failure rateMillions of disks at a few % AFR: one dies every few minutes. Eleven nines — and why copies alone are too costly.~7 min
Name & shape
The flat keyspace, its index, and immutable objects.
- 03Folders are a lie — the flat keyspace and its indexKeys are flat strings; the index maps key→location and shards by range, with a per-prefix throughput ceiling.~7 min
- 04Immutable objects: overwrite is a new version, not an edit — delete markers and the atomic pointer flipBytes never change: overwrite = new version + atomic pointer flip; delete drops a marker the old bytes hide behind.~7 min
Keep it alive
Erasure coding, placement, the two planes, and repair.
- 05Replication vs erasure coding: same safety, a third the costSplit into k+m fragments — any k rebuild it — tolerating m losses at ~1.4× instead of 3× replication.~7 min
- 05aAn m-tolerant code dies if m+1 fragments share a rackThe code's m-fault tolerance is real only if placement spreads fragments across independent failure domains.~7 min
- 06Two planes: the index says where, storage holds whatIndex plane vs data plane; trace a GET (read k, reconstruct, verify, stream) and a PUT's atomic index commit.~7 min
- 07The payoff: durability is repair winning a race — scrub, MTTR, and rebuilding lost fragmentsEleven nines is a rate equation: scrub finds rot, repair rebuilds lost fragments faster than disks destroy them.~7 min
Make it real
Strong consistency, multipart upload, lifecycle & tiering.
- 08Why strong consistency took S3 fourteen yearsEventual-consistency anomalies, the Dec 2020 flip, and the witness read-barrier — strong within a region only.~7 min
- 09Multipart upload and the ETag-that-isn't-an-MD5Parallel resumable parts, a completion manifest, and why the object ETag is a hash-of-hashes ending in -N.~7 min
- 10Lifecycle and tiering: cheap because bytes are immutableStorage classes trade retrieval latency for cost; a lifecycle rule slides an object down the ladder as a pointer move.~7 min
Design canvas
Assemble it for a workload and defend the durability number.
More in Object Storage, File Sync & Media Delivery
Durable bytes at rest and bytes in flight: an S3-style object store, a Dropbox-style sync client, and adaptive video streaming.
Prefer to design it yourself?
The same subject as a staged workspace: draw the architecture, and a simulator traces requests through the boxes you drew.
Open the Build an S3-style distributed object store workspace