Build a gRPC-style RPC framework
14 scenes · ~98 min · build the primitive

Build your own gRPC-style RPC framework

Every microservice talks over RPC, and the framework you ship determines half the system's failure modes. Build an RPC framework with codec, streams, deadlines, cancellation, retries, interceptors, and load-aware client-side balancing — and feel why gRPC ate the polyglot RPC market and why Thrift and JSON-over-HTTP linger.

Scenes
14 interactive scenes
Time
about 98 minutes
Topic
Caching, Proxies & the Edge

What you are building, and why

You have written a Flask, Express, or Spring service. You have called another service over HTTP with a client library, and you have used a TCP socket at least once. That is your starting point. By the end of this curriculum you will have built — conceptually, layer by layer — a gRPC-style RPC framework, and you will be able to design an RPC interface for a concrete workload: pick a codec, pick a transport, set a deadline and a retry policy, choose how the client balances across backends, choose an mTLS posture, and predict what breaks under packet loss or a brownout — and defend every choice with a named scene.

We start from the wire, not from the framework. Scene 1 is a single uncomfortable fact: a remote call is dressed up to look like greet("Ada"), but unlike a local call it can fail after the server already did the work — so the caller can be left not knowing whether it happened. Every later feature in the stack exists to manage one of the four ways that illusion leaks: latency, no shared memory, partial failure, and unordered concurrency. We name that frame once and call back to it the whole way up.

This is a curriculum about a specific density-of-vocabulary discipline: across 14 scenes there are exactly 26 named technical terms, and each scene introduces at most two. The diagram carries the load; the vocabulary follows. One running example — the call greet("Ada") — is threaded through every layer: you watch it become protobuf bytes, ride inside a DATA frame on one of many streams, travel under a shrinking deadline header, and finally prove its own identity over mTLS.

Resist the urge to "describe gRPC." Build each layer because the previous layer forced it, feel the failure mode each one prevents, and let the design canvas at the end push back on your choices.

What you will be able to explain afterwards

  • the remote-call illusion: it looks like a local function but can fail after the work is done (partial failure)
  • framing on a TCP byte stream: recv() is not a message; length-prefix to find the boundaries
  • IDL + codec: protobuf field numbers (not names) as the schema-evolution superpower; never reuse a number
  • HTTP/2 as transport: one RPC = one stream; many streams multiplexed over one connection
  • unary, server-stream, client-stream, bidi — four shapes from one stream of DATA frames
  • deadlines, not timeouts: an absolute clock that PROPAGATES across an N-hop chain (the single most important feature)
  • cancellation as a first-class event (RST_STREAM); the Context object carries deadline + cancel together
  • retries: idempotency gates whether, a token-bucket budget + jittered backoff gate how much (retry storms / metastable failure)
  • interceptors / middleware: auth, metrics, tracing, retry composed as one onion around every call
  • client-side load balancing + discovery: why an L4 LB pins every stream to one backend (no central LB)
  • flow control / backpressure via the HTTP/2 window; a slow reader slows the writer
  • HTTP/2 fixed app-layer head-of-line blocking; TCP-layer HOL remains; QUIC/HTTP3 fixes it
  • TLS vs mTLS: channel encryption vs verified workload identity, propagated across the call chain
  1. 01
  2. 02
  3. 03
  4. 04
  5. 05
  6. 06
  7. 06a
  8. 07
  9. 08
  10. 09
  11. 10
  12. 10a
  13. 11
  14. 12

The wire

Bytes, frames, and the schema that gives them meaning.

  1. 01
    Calling a function on another machine — partial failure behind the RPC stub
    A remote call is dressed up to look like a local one — but unlike a local call it can fail after the server already did the work, so the caller can't tell if it happened.
    ~7 min
  2. 02
    TCP gives you bytes, not messages
    TCP is a byte stream, not a message stream: two messages can arrive glued or split. A length prefix is what lets the receiver read exactly one message back.
    ~7 min
  3. 03
    The schema: field numbers, not field names
    A schema keys each field by a stable number, not its name — so an old reader can skip a field it doesn't know and still decode the rest. Never reuse a field number.
    ~7 min

The transport

HTTP/2 streams and the four shapes of a call.

  1. 04
    One RPC, one stream, one pipe for many
    HTTP/2 multiplexes many independent streams over one long-lived TCP connection; gRPC maps one RPC to one stream. A hundred calls share one pipe, not a hundred sockets.
    ~7 min
  2. 05
    Four call shapes from one stream — half-close and DATA frames on HTTP/2
    Unary and the three streaming shapes are the same stream — only the number and direction of DATA frames differ. Streaming is a consequence of the transport, not a bolt-on.
    ~7 min

Reliability

Deadlines, cancellation, retries, and the interceptor onion.

  1. 06
    Deadlines, not timeouts — deadline propagation with the grpc-timeout header
    A timeout is relative and resets each hop; a deadline is absolute and propagates the time remaining — so a downstream service never works for a caller that already gave up.
    ~7 min
  2. 06a
    Cancellation: an event, not a clock — gRPC Context cancellation and RST_STREAM
    A deadline fires from a clock; cancellation fires from an event. Both ride one Context object down the chain, so aborting a parent stops all the doomed downstream work.
    ~7 min
  3. 07
    Retries: idempotency and a token budget
    Only an idempotent method is safe to auto-retry, and even then a token-bucket budget must cap retries — or a brownout turns into a self-sustaining retry storm.
    ~7 min
  4. 08
    Interceptors: the middleware onion
    An interceptor wraps every call as one composable layer, so auth, metrics, tracing, and the retry policy are written once around the handler instead of per method.
    ~7 min

Resilience

Client-side balancing, backpressure, head-of-line blocking.

  1. 09
    The L4 pinning trap: balance requests, not connections
    HTTP/2's one long-lived connection means an L4 load balancer pins every RPC to one backend. Client-side balancing discovers all backends and picks one per request.
    ~7 min
  2. 10
    Flow control: a slow reader slows the writer
    The receiver advertises a window of credit; the sender may only send DATA up to it. A slow reader stops granting credit, so the producer pauses instead of OOMing.
    ~7 min
  3. 10a
    Head-of-line blocking and the QUIC fix
    HTTP/2 fixed app-layer head-of-line blocking, but all streams share one in-order TCP pipe, so one lost packet stalls them all. QUIC moves streams below the loss boundary.
    ~7 min

Secure & ship

mTLS identity, then configure the whole stack.

  1. 11
    mTLS and propagating identity
    TLS encrypts and proves the server; mTLS proves both peers with a workload identity. The original caller's identity must propagate across hops, just like a deadline.
    ~7 min
  2. 12
    Design canvas: configure the RPC stack
    Pick codec, transport, deadline, retry, balancing, and security for four named workloads — each defended by the scene that taught it. gRPC inside the fleet, REST at the edge.
    ~7 min

Prefer to design it yourself?

The same subject as a staged workspace: draw the architecture, and a simulator traces requests through the boxes you drew.

Open the Build a gRPC-style RPC framework workspace