Build a distributed tracing system (Jaeger / Zipkin style)
17 scenes · ~119 min · build the primitive

Build your own distributed tracing system (Jaeger / Zipkin style)

Metrics tell you that checkout is slow and logs tell you what one service printed; neither can say which hop of one request consumed the time. Build the third pillar: one record per hop tied together by one id — span propagation, sampling that doesn't lie, an export path that never blocks a customer, storage keyed by trace id, and the waterfall that finally answers 'where did the latency go?'

Scenes
17 interactive scenes
Time
about 119 minutes
Topic
Observability: Metrics, Logs & Traces

What you are building, and why

You are building the tracing system for Shopfront, the same online store whose metrics platform you built in metrics-system. A checkout that should take 200 ms is taking 3 seconds. Five teams are on the call. Every one of them has a dashboard, every dashboard is green, and nobody can prove which hop of that one request burned the time.

That is not a gap in anyone's diligence. Logs and metrics are both recorded per service, so neither can answer a question about one request across services. Fan-out amplifies the tail multiplicatively — Dean & Barroso's canonical example is 100 servers at a 1-in-100 chance of a one-second response, which puts 63% of requests over a second even though every individual server looks fine. A service's own p99 is computed over its own calls, not over the requests that touched it. And time spent queued between services appears in nobody's self-reported latency at all.

This problem builds the third pillar: one record per hop of one request, all stamped with the same id, so the slow hop names itself instead of being argued about. metrics-system gets someone paged; tracing gets them to the culprit.

One sentence runs through every stage, and it is what makes tracing harder than it looks: an incomplete trace does not look incomplete. A dropped header, an independent sampling decision, a full export queue, a late span and a skewed clock all produce a picture that is wrong and still looks finished. Nothing throws. There is no error state. The trace just quietly means something else than what it appears to mean — and the whole design exists to make that failure mode visible, bounded, or impossible.

Keep one number honest the whole way: time to culprit, the minutes between the page firing and an engineer naming the service and operation responsible.

What you will be able to explain afterwards

  • spans, traces, and the parent-child relationship rebuilt at read time
  • context propagation: W3C Trace Context, B3, and the async boundary
  • why a queue hop drops the chain, and links vs parents for batches
  • high cardinality: cheap on a span, ruinous as a metric label
  • head sampling: one decision per trace, p^N, and the sampled flag
  • sampling bias and adjusted counts (weight by the inverse of the keep rate)
  • tail sampling: trace-id routing, a decision window, and late spans
  • bounded lossy export: drop telemetry, never block the request path
  • the collector tier: enrich, redact, fan out, and where tail decisions must live
  • trace storage: random trace-id partitioning, and why a trace is never closed
  • finding a trace: the index tax, the scan tax, and arriving by exemplar
  • reading a waterfall: self time, the critical path, and what a gap means
  • clock skew, the viewer's adjustment, and what that adjustment hides
  • derived signals: span metrics and service graphs, and how sampling distorts them
  1. 01
  2. 02
  3. 03
  4. 04
  5. 05
  6. 06
  7. 07
  8. 08
  9. 09
  10. 10
  11. 11
  12. 12
  13. 13
  14. 14
  15. 14a
  16. 15
  17. 16

Why trace

Per-service records can't blame a hop; one id can.

  1. 01
    Four teams, one slow checkout — traces, spans and time-to-culprit
    Logs and dashboards are both recorded per service, so neither can name the hop that consumed the time. One record per hop of a single request, tied together, makes the slow hop name itself.
    ~7 min
  2. 02
    One row per hop, and the id that links them — parent span ids and read-time assembly
    Each hop writes one flat record carrying its own id, its parent's id and a shared trace id. Nothing is nested when it is written — the tree is rebuilt later, which is why a missing record leaves a plausible-looking hole.
    ~7 min

Carry it

The header, the queue that drops it, what rides along.

  1. 03
    The id has to ride along — context propagation with the traceparent header
    The caller writes the trace id and its own span id into an outbound header and the callee reads them back. Any hop that builds a request without copying it starts a brand-new trace — and both halves still look complete.
    ~7 min
  2. 04
    Where the chain breaks: queues, pools, batches — context injection and span links across queues
    A queue hop has no request to attach a header to, so the context must ride inside the message. And because one batch can have many causes while a span has exactly one parent, the consumer records a link to each.
    ~7 min
  3. 05
    High cardinality is free here
    A span is already one row per request per hop, so attaching a customer id adds bytes to a row that was being written anyway and multiplies nothing — the exact opposite of a metric label.
    ~7 min

Sampling

You can't keep it all; decide once, and don't lie.

  1. 06
    Every request, every hop, forever? — full capture cost and the keep rate
    Spans multiply: requests per second times hops per request. Full capture is a storage bill plus a measured latency cost on the request path, which makes traces the one pillar where you deliberately keep a fraction.
    ~7 min
  2. 07
    Decide once, carry the answer — head-based sampling and the sampled flag
    Sampling is per trace, not per span. If N hops each flip their own p coin a whole trace survives with probability p to the N, and what you store is mostly fragments that look like whole traces.
    ~7 min
  3. 08
    Sampling that doesn't lie — sampling bias and adjusted counts
    A uniform sample keeps proportions but destroys rare events; rate limiting over-represents quiet services; and counting stored traces without weighting each by the inverse of its keep rate understates everything.
    ~7 min
  4. 09
    Decide after you've seen the whole trace — tail-based sampling and the decision window
    To judge a trace on whole-trace facts you must route every span of it to the same decision point, hold them for a window, and accept that anything arriving later is judged on its own — the queue hop again.
    ~7 min

Pipeline

Off the hot path, then a tier that isn't your app.

  1. 10
    Never make a customer wait for telemetry
    A finished span goes into a bounded queue that a background worker drains in batches. When it fills, the library drops spans and counts them — blocking the request would turn a backend slowdown into an outage.
    ~7 min
  2. 11
    A layer that isn't your app — collector agents and the gateway pool
    A collector tier takes the jobs no application should do: stamping where a span ran, stripping anything sensitive, absorbing bursts, fanning out. Tail decisions live in its pool, because only a pool can be routed to.
    ~7 min

Store & read

Key by trace id, find it, read it, distrust it.

  1. 12
    The write path: one row per span, one key per trace
    Spans are written as they arrive under a random trace id that spreads perfectly across partitions. Nothing ever closes a trace, so a partial trace is the normal steady state and the tree is rebuilt at read time.
    ~7 min
  2. 13
    Nobody knows the trace id
    Every way into the store other than the id costs something: an index prices high-cardinality attributes, a scan trades that storage for query compute. The cheapest way in is to arrive already holding the id.
    ~7 min
  3. 14
    Where did the latency go? — critical path and self time in a waterfall
    The waterfall answers three questions at a glance: which span holds time its children don't account for, which children overlap, and where there is no span at all. Only the chain that doesn't overlap can be shortened.
    ~7 min
  4. 14a
    When the picture itself lies — clock skew, gaps and the uninstrumented hop
    Every span's times come from the clock of the machine that recorded it, and those clocks disagree. Viewers quietly shift timestamps to make a child fit its parent, folding real network and queue wait into the parent.
    ~7 min

Ship it

What a million traces say, then design the stack.

  1. 15
    What a million traces say together — the service dependency graph and span metrics
    The parent-child edges of many traces draw the live dependency graph, and counting spans by service and operation yields rates and error ratios for hops no dashboard covered — both only as honest as the sampling.
    ~7 min
  2. 16
    Design your tracing stack
    A tracing stack is a chain of decisions: what to instrument, what to propagate, how much to keep and by what rule, where that decision is made, and how a trace gets found. Each is right only for a given workload.
    ~7 min

Prefer to design it yourself?

The same subject as a staged workspace: draw the architecture, and a simulator traces requests through the boxes you drew.

Open the Build a distributed tracing system (Jaeger / Zipkin style) workspace