Build your own Service Mesh (Envoy / Istio style)
Every microservice request crosses two proxies. This curriculum is what they do: routing, load balancing, timeout-and-retry-budget, circuit breakers, outlier detection, token-bucket rate limits, mTLS with workload identity, and a control plane that streams config to all of them. Build it in the order the production problems show up — and feel why Envoy plus a control plane has eaten the east-west world.
- Scenes
- 13 interactive scenes
- Time
- about 91 minutes
- Topic
- Caching, Proxies & the Edge
What you are building, and why
You have written a Flask, Express, or Spring service. You have called another service over HTTP. You have used curl -v and watched the TLS handshake happen. That is your starting point. By the end of this curriculum you will be able to design a service-mesh deployment — pick sidecar vs edge proxy, route on path or header, set timeout / retry budget / circuit breaker / outlier detection thresholds for a concrete workload, choose mTLS posture, decide local vs global rate limit, and predict what happens when the control plane dies — and defend each choice with a named scene.
This curriculum is the working developer's guide to the piece of infrastructure every microservice request passes through, in the order the production problems show up. It is also a curriculum about a specific, density-of-vocabulary realization: across 13 scenes there are roughly 28 named technical terms, and each scene introduces at most two. The visual carries the load; the vocabulary follows.
Resist the urge to "describe Envoy" or "describe Istio." Make decisions yourself, defend them, and let the design push back. Every knob you can name on the design canvas is a knob you can defend in a production review.
What you will be able to explain afterwards
- the 50-services problem: heterogeneous policy → cascading outage
- sidecar pattern; data plane vs control plane
- L4 vs L7 proxying (which sees path/headers, which doesn't)
- listener + ordered route table; canary by header or weight
- cluster as group-of-replicas; load-balancing policies (round-robin / least-request / ring-hash)
- timeout, per-try timeout, exponential backoff with jitter
- retry budget — the fix for retry storms (≤3% Envoy default)
- circuit breaker — closed / open / half-open
- outlier detection (passive eject) vs active health check
- token-bucket rate limiting; local (cheap, drifts) vs global (exact, RPC hop)
- mTLS, SPIFFE-style workload identity, short-lived certs
- control plane vs data plane; xDS (LDS / RDS / CDS / EDS / SDS) streaming
- trace + span; W3C traceparent propagation; sampling vs RED metrics
- design canvas: pick posture per workload (latency-critical, ingress, batch, partner)
- 01Fifty services, fifty broken retry policiesEvery team picks its own retry, timeout, breaker, and mTLS library. One slow dependency turns into a fleet-wide outage.~7 min
- 02The sidecar — one proxy per podPut a small proxy next to every service. App talks to localhost; cross-service traffic flows sidecar to sidecar. The fleet of sidecars is the service mesh.~7 min
- 03L4 vs L7 — bytes or requestsAn L4 proxy forwards opaque TCP bytes; an L7 proxy parses HTTP and can act on path, method, and headers. The mesh is L7 for everything that follows.~7 min
- 04Listener and route — bind and matchA listener accepts on a port; an ordered route table picks a destination on the first match. Reorder the rules and the same request lands somewhere else.~7 min
- 05Cluster and load balancing — pick one of manyBehind one destination name is a cluster of replicas; round-robin, least-request, or ring-hash decides who serves each request. Round-robin is the wrong default under heterogeneous latency.~7 min
- 06Timeout and retry budget — bounded patienceNaive multi-hop retries amplify load 243x on a failing backend. A retry budget caps total retries as a fraction of normal traffic so retries can't become the outage.~7 min
- 07Circuit breaker — the state machineClosed → open → half-open. Fast-fail to a known-broken dependency and periodically probe for recovery, so callers stop wasting resources on guaranteed failures.~7 min
- 07aOutlier detection — eject one bad replicaDon't trip the whole cluster — pull just the misbehaving replica from the pool. Passive (real 5xx) catches what active /healthz probes miss.~7 min
- 08Rate limiting — the token bucketPer-client token bucket: each request takes a token, an empty bucket returns 429. Local is cheap and drifts; global stays exact via a coordinator.~7 min
- 09mTLS — identity for both sidesTLS proves the server, mTLS proves both. The control plane mints short-lived certs that carry a stable workload identity — pod IPs are not identities.~7 min
- 10Control plane and data plane — config over gRPCSidecars (data plane) handle traffic; the control plane (Istiod) streams listener/route/cluster/cert config via xDS. Kill the control plane and traffic keeps flowing.~7 min
- 11Trace and span — stitching one user requestEvery sidecar emits a span tagged with the trace id from `traceparent`. One forgotten header rebuild breaks the trace silently — RED metrics keep flowing regardless.~7 min
- 12Design canvas — configure the meshFour workloads, every knob from the prior scenes. Each verifier note cites the scene that earned it.~7 min
More in Caching, Proxies & the Edge
Everything between the client and the origin: in-memory caches, CDNs, load balancers and service proxies — and the three ways a cache betrays you.
- Cache Invalidation Across a FleetWrite-through vs write-behind. Two generals.
- Build Build RedisAn in-memory data-structure server: one thread, rich types, optional persistence, async replication. Internalize the cost of single-threaded simplicity and a dozen caching/HA decisions get easier.
- Build Build a CDNA globally-distributed reverse proxy whose only job is to (a) terminate the user's TCP/TLS milliseconds away and (b) serve a cached origin response so origin never sees the request. Internalize edge caching, anycast, TTL, revalidation, SWR, purge, the Vary footgun, origin shield, bypass, and hit ratio — and the dozen ways to misconfigure each.
- Build Build a gRPC-style RPC frameworkEvery microservice talks over RPC, and the framework you ship determines half the system's failure modes. Build an RPC framework with codec, streams, deadlines, cancellation, retries, interceptors, and load-aware client-side balancing — and feel why gRPC ate the polyglot RPC market and why Thrift and JSON-over-HTTP linger.
Prefer to design it yourself?
The same subject as a staged workspace: draw the architecture, and a simulator traces requests through the boxes you drew.
Open the Build a Service Mesh (Envoy / Istio style) workspace