Build your own gRPC-style RPC framework
Every microservice talks over RPC, and the framework you ship determines half the system's failure modes. Build an RPC framework with codec, streams, deadlines, cancellation, retries, interceptors, and load-aware client-side balancing — and feel why gRPC ate the polyglot RPC market and why Thrift and JSON-over-HTTP linger.
- Scenes
- 14 interactive scenes
- Time
- about 98 minutes
- Topic
- Caching, Proxies & the Edge
What you are building, and why
You have written a Flask, Express, or Spring service. You have called another service over HTTP with a client library, and you have used a TCP socket at least once. That is your starting point. By the end of this curriculum you will have built — conceptually, layer by layer — a gRPC-style RPC framework, and you will be able to design an RPC interface for a concrete workload: pick a codec, pick a transport, set a deadline and a retry policy, choose how the client balances across backends, choose an mTLS posture, and predict what breaks under packet loss or a brownout — and defend every choice with a named scene.
We start from the wire, not from the framework. Scene 1 is a single uncomfortable fact: a remote call is dressed up to look like greet("Ada"), but unlike a local call it can fail after the server already did the work — so the caller can be left not knowing whether it happened. Every later feature in the stack exists to manage one of the four ways that illusion leaks: latency, no shared memory, partial failure, and unordered concurrency. We name that frame once and call back to it the whole way up.
This is a curriculum about a specific density-of-vocabulary discipline: across 14 scenes there are exactly 26 named technical terms, and each scene introduces at most two. The diagram carries the load; the vocabulary follows. One running example — the call greet("Ada") — is threaded through every layer: you watch it become protobuf bytes, ride inside a DATA frame on one of many streams, travel under a shrinking deadline header, and finally prove its own identity over mTLS.
Resist the urge to "describe gRPC." Build each layer because the previous layer forced it, feel the failure mode each one prevents, and let the design canvas at the end push back on your choices.
What you will be able to explain afterwards
- the remote-call illusion: it looks like a local function but can fail after the work is done (partial failure)
- framing on a TCP byte stream: recv() is not a message; length-prefix to find the boundaries
- IDL + codec: protobuf field numbers (not names) as the schema-evolution superpower; never reuse a number
- HTTP/2 as transport: one RPC = one stream; many streams multiplexed over one connection
- unary, server-stream, client-stream, bidi — four shapes from one stream of DATA frames
- deadlines, not timeouts: an absolute clock that PROPAGATES across an N-hop chain (the single most important feature)
- cancellation as a first-class event (RST_STREAM); the Context object carries deadline + cancel together
- retries: idempotency gates whether, a token-bucket budget + jittered backoff gate how much (retry storms / metastable failure)
- interceptors / middleware: auth, metrics, tracing, retry composed as one onion around every call
- client-side load balancing + discovery: why an L4 LB pins every stream to one backend (no central LB)
- flow control / backpressure via the HTTP/2 window; a slow reader slows the writer
- HTTP/2 fixed app-layer head-of-line blocking; TCP-layer HOL remains; QUIC/HTTP3 fixes it
- TLS vs mTLS: channel encryption vs verified workload identity, propagated across the call chain
The wire
Bytes, frames, and the schema that gives them meaning.
- 01Calling a function on another machine — partial failure behind the RPC stubA remote call is dressed up to look like a local one — but unlike a local call it can fail after the server already did the work, so the caller can't tell if it happened.~7 min
- 02TCP gives you bytes, not messagesTCP is a byte stream, not a message stream: two messages can arrive glued or split. A length prefix is what lets the receiver read exactly one message back.~7 min
- 03The schema: field numbers, not field namesA schema keys each field by a stable number, not its name — so an old reader can skip a field it doesn't know and still decode the rest. Never reuse a field number.~7 min
The transport
HTTP/2 streams and the four shapes of a call.
- 04One RPC, one stream, one pipe for manyHTTP/2 multiplexes many independent streams over one long-lived TCP connection; gRPC maps one RPC to one stream. A hundred calls share one pipe, not a hundred sockets.~7 min
- 05Four call shapes from one stream — half-close and DATA frames on HTTP/2Unary and the three streaming shapes are the same stream — only the number and direction of DATA frames differ. Streaming is a consequence of the transport, not a bolt-on.~7 min
Reliability
Deadlines, cancellation, retries, and the interceptor onion.
- 06Deadlines, not timeouts — deadline propagation with the grpc-timeout headerA timeout is relative and resets each hop; a deadline is absolute and propagates the time remaining — so a downstream service never works for a caller that already gave up.~7 min
- 06aCancellation: an event, not a clock — gRPC Context cancellation and RST_STREAMA deadline fires from a clock; cancellation fires from an event. Both ride one Context object down the chain, so aborting a parent stops all the doomed downstream work.~7 min
- 07Retries: idempotency and a token budgetOnly an idempotent method is safe to auto-retry, and even then a token-bucket budget must cap retries — or a brownout turns into a self-sustaining retry storm.~7 min
- 08Interceptors: the middleware onionAn interceptor wraps every call as one composable layer, so auth, metrics, tracing, and the retry policy are written once around the handler instead of per method.~7 min
Resilience
Client-side balancing, backpressure, head-of-line blocking.
- 09The L4 pinning trap: balance requests, not connectionsHTTP/2's one long-lived connection means an L4 load balancer pins every RPC to one backend. Client-side balancing discovers all backends and picks one per request.~7 min
- 10Flow control: a slow reader slows the writerThe receiver advertises a window of credit; the sender may only send DATA up to it. A slow reader stops granting credit, so the producer pauses instead of OOMing.~7 min
- 10aHead-of-line blocking and the QUIC fixHTTP/2 fixed app-layer head-of-line blocking, but all streams share one in-order TCP pipe, so one lost packet stalls them all. QUIC moves streams below the loss boundary.~7 min
Secure & ship
mTLS identity, then configure the whole stack.
- 11mTLS and propagating identityTLS encrypts and proves the server; mTLS proves both peers with a workload identity. The original caller's identity must propagate across hops, just like a deadline.~7 min
- 12Design canvas: configure the RPC stackPick codec, transport, deadline, retry, balancing, and security for four named workloads — each defended by the scene that taught it. gRPC inside the fleet, REST at the edge.~7 min
More in Caching, Proxies & the Edge
Everything between the client and the origin: in-memory caches, CDNs, load balancers and service proxies — and the three ways a cache betrays you.
- Cache Invalidation Across a FleetWrite-through vs write-behind. Two generals.
- Build Build RedisAn in-memory data-structure server: one thread, rich types, optional persistence, async replication. Internalize the cost of single-threaded simplicity and a dozen caching/HA decisions get easier.
- Build Build a CDNA globally-distributed reverse proxy whose only job is to (a) terminate the user's TCP/TLS milliseconds away and (b) serve a cached origin response so origin never sees the request. Internalize edge caching, anycast, TTL, revalidation, SWR, purge, the Vary footgun, origin shield, bypass, and hit ratio — and the dozen ways to misconfigure each.
- Build Build a Service Mesh (Envoy / Istio style)Every microservice request crosses two proxies. This curriculum is what they do: routing, load balancing, timeout-and-retry-budget, circuit breakers, outlier detection, token-bucket rate limits, mTLS with workload identity, and a control plane that streams config to all of them. Build it in the order the production problems show up — and feel why Envoy plus a control plane has eaten the east-west world.
Prefer to design it yourself?
The same subject as a staged workspace: draw the architecture, and a simulator traces requests through the boxes you drew.
Open the Build a gRPC-style RPC framework workspace