Build a distributed cache (Memcached / Pelikan style)
One Redis is easy. A hundred Redises serving 10M QPS with sub-millisecond p99 across a fleet is the real test. Build the client-side consistent-hashing distributed cache: hot-key replication, mirroring for fault tolerance, the thundering-herd dogpile, and the cache-coherence gymnastics that come with sharding.Enter to send · Shift+Enter for a new line
About Build a distributed cache (Memcached / Pelikan style)
One Redis is easy. A hundred Redises serving 10M QPS with sub-millisecond p99 across a fleet is the real test. Build the client-side consistent-hashing distributed cache: hot-key replication, mirroring for fault tolerance, the thundering-herd dogpile, and the cache-coherence gymnastics that come with sharding.
- Difficulty
- intermediate
- Time
- about 75 minutes
- Stages
- 9
- Topic
- Caching, Proxies & the Edge
How this problem is worked
Nine stages, from what the thing is for to how it compares with the real implementations. Each asks one question, and the simulator runs the architecture you draw against the requirements you wrote.
- 01Purpose & invariantsWhat is this for, and what must always be true of it?
- 02Workload characterizationWho writes, who reads, and in what shapes?
- 03Data model & on-disk formatWhat does the data look like at rest?
- 04Core algorithmsHow do the write path and the read path actually work?
- 05Distribution & replicationHow does this scale out and survive losing a machine?
- 06Consistency & correctnessUnder concurrency and failure, what is guaranteed?
- 07Failure modes & recoveryWhat actually happens when each part fails?
- 08Operational characteristicsCan a human run this at three in the morning?
- 09Trade-offs & comparisonWhere does this sit against the alternatives?
Primary sources for this problem
- Nishtala et al. — Scaling Memcache at Facebook (NSDI 2013)
- Memcached protocol + Twitter Twemcache / Pelikan engineering blog
- Twitter — Twemproxy (nutcracker) README and design
- Facebook — mcrouter design doc
- Karger et al. — Consistent Hashing and Random Trees (STOC 1997)
- Cliff Click — A Lock-Free Hash Table (the per-server side)
More in Caching, Proxies & the Edge
Everything between the client and the origin: in-memory caches, CDNs, load balancers and service proxies — and the three ways a cache betrays you.
- Cache Invalidation Across a FleetWrite-through vs write-behind. Two generals.
- Build Build RedisAn in-memory data-structure server: one thread, rich types, optional persistence, async replication. Internalize the cost of single-threaded simplicity and a dozen caching/HA decisions get easier.
- Build Build a CDNA globally-distributed reverse proxy whose only job is to (a) terminate the user's TCP/TLS milliseconds away and (b) serve a cached origin response so origin never sees the request. Internalize edge caching, anycast, TTL, revalidation, SWR, purge, the Vary footgun, origin shield, bypass, and hit ratio — and the dozen ways to misconfigure each.
- Build Build a Service Mesh (Envoy / Istio style)Every microservice request crosses two proxies. This curriculum is what they do: routing, load balancing, timeout-and-retry-budget, circuit breakers, outlier detection, token-bucket rate limits, mTLS with workload identity, and a control plane that streams config to all of them. Build it in the order the production problems show up — and feel why Envoy plus a control plane has eaten the east-west world.
- Build Build a gRPC-style RPC frameworkEvery microservice talks over RPC, and the framework you ship determines half the system's failure modes. Build an RPC framework with codec, streams, deadlines, cancellation, retries, interceptors, and load-aware client-side balancing — and feel why gRPC ate the polyglot RPC market and why Thrift and JSON-over-HTTP linger.
Browse the full problem catalog, or see what the simulator does and does not model.