Build a document store (MongoDB-style)
Schemaless documents that still behave like a database. Build a system that stores JSON/BSON natively, ships replica sets with primary elections, shards a collection by key range or hash, and keeps secondary indexes consistent under write — and feel exactly where 'flexible schema' turns into 'silent inconsistency'.Enter to send · Shift+Enter for a new line
About Build a document store (MongoDB-style)
Schemaless documents that still behave like a database. Build a system that stores JSON/BSON natively, ships replica sets with primary elections, shards a collection by key range or hash, and keeps secondary indexes consistent under write — and feel exactly where 'flexible schema' turns into 'silent inconsistency'.
- Difficulty
- intermediate
- Time
- about 85 minutes
- Stages
- 9
- Topic
- Storage Engines & Databases
How this problem is worked
Nine stages, from what the thing is for to how it compares with the real implementations. Each asks one question, and the simulator runs the architecture you draw against the requirements you wrote.
- 01Purpose & invariantsWhat is this for, and what must always be true of it?
- 02Workload characterizationWho writes, who reads, and in what shapes?
- 03Data model & on-disk formatWhat does the data look like at rest?
- 04Core algorithmsHow do the write path and the read path actually work?
- 05Distribution & replicationHow does this scale out and survive losing a machine?
- 06Consistency & correctnessUnder concurrency and failure, what is guaranteed?
- 07Failure modes & recoveryWhat actually happens when each part fails?
- 08Operational characteristicsCan a human run this at three in the morning?
- 09Trade-offs & comparisonWhere does this sit against the alternatives?
Primary sources for this problem
- MongoDB Architecture Guide (mongodb.com whitepaper)
- MongoDB Manual — Replica sets, Sharding, Read/Write concerns
- WiredTiger architecture doc (storage engine internals)
- MongoDB blog — Causal consistency in MongoDB 3.6+
- Kleppmann — Designing Data-Intensive Applications (Ch. 2 — document vs relational)
- Jepsen — MongoDB analyses (multiple)
More in Storage Engines & Databases
Open the box every design diagram labels "DB": pages, logs, LSM trees, wide-column, documents, graphs and columnar scans, built from scratch.
- Build Build a Bitcask-style KV storeThe simplest possible KV store that still works: an append-only log on disk + an in-memory hash index. Build it from first principles and feel which trade-offs every later store inherits.
- Build Build an LSM-tree storage engine (LevelDB / RocksDB style)The simplest possible storage engine that gives you BOTH ordered reads AND more keys than fit in RAM, by accepting a deal: write to RAM at memory speed, log to disk for safety, then merge sorted files in the background forever.
- Build Build a B-tree storage engine (SQLite-style)What actually happens when you run INSERT INTO users(...). One file of fixed-size pages, organized as B-trees, with a write-ahead log that turns commits into appends. Build it from a SQL writer's perspective and feel why every knob exists.
- Build Build a wide-column store (Cassandra / DynamoDB family)One server is not enough — disk fills, throughput maxes, the box dies. Build a multi-node store from first principles: hash sharding, the consistent-hash ring, vnodes, replication, eventual consistency, tunable W+R quorum, hinted handoff, read repair. Every modern Dynamo-style store is a point in this design space.
- Build Build a graph database (Neo4j / Dgraph-style)When the workload is 'friends of friends', a relational join melts. Build a store where edges are first-class — index-free adjacency, traversals that follow pointers instead of joining tables, and a query language (Cypher / GraphQL+) that thinks in patterns. Feel why graph storage shines for traversal-heavy work and stumbles on full-graph aggregates.
- Build Build a columnar OLAP store (ClickHouse / Druid style)OLTP picks one row by key; OLAP scans a billion rows of one column and asks for a percentile. Build the analytical engine that makes that fast: columnar layout, dictionary/RLE/delta compression, vectorized execution, late materialization, MPP shuffle. Internalize why Postgres is 1000× slower than ClickHouse on the same query and why the inverse is also true.
Browse the full problem catalog, or see what the simulator does and does not model.