Build your own distributed tracing system (Jaeger / Zipkin style)
Metrics tell you that checkout is slow and logs tell you what one service printed; neither can say which hop of one request consumed the time. Build the third pillar: one record per hop tied together by one id — span propagation, sampling that doesn't lie, an export path that never blocks a customer, storage keyed by trace id, and the waterfall that finally answers 'where did the latency go?'
- Scenes
- 17 interactive scenes
- Time
- about 119 minutes
- Topic
- Observability: Metrics, Logs & Traces
What you are building, and why
You are building the tracing system for Shopfront, the same online store whose metrics platform you built in metrics-system. A checkout that should take 200 ms is taking 3 seconds. Five teams are on the call. Every one of them has a dashboard, every dashboard is green, and nobody can prove which hop of that one request burned the time.
That is not a gap in anyone's diligence. Logs and metrics are both recorded per service, so neither can answer a question about one request across services. Fan-out amplifies the tail multiplicatively — Dean & Barroso's canonical example is 100 servers at a 1-in-100 chance of a one-second response, which puts 63% of requests over a second even though every individual server looks fine. A service's own p99 is computed over its own calls, not over the requests that touched it. And time spent queued between services appears in nobody's self-reported latency at all.
This problem builds the third pillar: one record per hop of one request, all stamped with the same id, so the slow hop names itself instead of being argued about. metrics-system gets someone paged; tracing gets them to the culprit.
One sentence runs through every stage, and it is what makes tracing harder than it looks: an incomplete trace does not look incomplete. A dropped header, an independent sampling decision, a full export queue, a late span and a skewed clock all produce a picture that is wrong and still looks finished. Nothing throws. There is no error state. The trace just quietly means something else than what it appears to mean — and the whole design exists to make that failure mode visible, bounded, or impossible.
Keep one number honest the whole way: time to culprit, the minutes between the page firing and an engineer naming the service and operation responsible.
What you will be able to explain afterwards
- spans, traces, and the parent-child relationship rebuilt at read time
- context propagation: W3C Trace Context, B3, and the async boundary
- why a queue hop drops the chain, and links vs parents for batches
- high cardinality: cheap on a span, ruinous as a metric label
- head sampling: one decision per trace, p^N, and the sampled flag
- sampling bias and adjusted counts (weight by the inverse of the keep rate)
- tail sampling: trace-id routing, a decision window, and late spans
- bounded lossy export: drop telemetry, never block the request path
- the collector tier: enrich, redact, fan out, and where tail decisions must live
- trace storage: random trace-id partitioning, and why a trace is never closed
- finding a trace: the index tax, the scan tax, and arriving by exemplar
- reading a waterfall: self time, the critical path, and what a gap means
- clock skew, the viewer's adjustment, and what that adjustment hides
- derived signals: span metrics and service graphs, and how sampling distorts them
Why trace
Per-service records can't blame a hop; one id can.
- 01Four teams, one slow checkout — traces, spans and time-to-culpritLogs and dashboards are both recorded per service, so neither can name the hop that consumed the time. One record per hop of a single request, tied together, makes the slow hop name itself.~7 min
- 02One row per hop, and the id that links them — parent span ids and read-time assemblyEach hop writes one flat record carrying its own id, its parent's id and a shared trace id. Nothing is nested when it is written — the tree is rebuilt later, which is why a missing record leaves a plausible-looking hole.~7 min
Carry it
The header, the queue that drops it, what rides along.
- 03The id has to ride along — context propagation with the traceparent headerThe caller writes the trace id and its own span id into an outbound header and the callee reads them back. Any hop that builds a request without copying it starts a brand-new trace — and both halves still look complete.~7 min
- 04Where the chain breaks: queues, pools, batches — context injection and span links across queuesA queue hop has no request to attach a header to, so the context must ride inside the message. And because one batch can have many causes while a span has exactly one parent, the consumer records a link to each.~7 min
- 05High cardinality is free hereA span is already one row per request per hop, so attaching a customer id adds bytes to a row that was being written anyway and multiplies nothing — the exact opposite of a metric label.~7 min
Sampling
You can't keep it all; decide once, and don't lie.
- 06Every request, every hop, forever? — full capture cost and the keep rateSpans multiply: requests per second times hops per request. Full capture is a storage bill plus a measured latency cost on the request path, which makes traces the one pillar where you deliberately keep a fraction.~7 min
- 07Decide once, carry the answer — head-based sampling and the sampled flagSampling is per trace, not per span. If N hops each flip their own p coin a whole trace survives with probability p to the N, and what you store is mostly fragments that look like whole traces.~7 min
- 08Sampling that doesn't lie — sampling bias and adjusted countsA uniform sample keeps proportions but destroys rare events; rate limiting over-represents quiet services; and counting stored traces without weighting each by the inverse of its keep rate understates everything.~7 min
- 09Decide after you've seen the whole trace — tail-based sampling and the decision windowTo judge a trace on whole-trace facts you must route every span of it to the same decision point, hold them for a window, and accept that anything arriving later is judged on its own — the queue hop again.~7 min
Pipeline
Off the hot path, then a tier that isn't your app.
- 10Never make a customer wait for telemetryA finished span goes into a bounded queue that a background worker drains in batches. When it fills, the library drops spans and counts them — blocking the request would turn a backend slowdown into an outage.~7 min
- 11A layer that isn't your app — collector agents and the gateway poolA collector tier takes the jobs no application should do: stamping where a span ran, stripping anything sensitive, absorbing bursts, fanning out. Tail decisions live in its pool, because only a pool can be routed to.~7 min
Store & read
Key by trace id, find it, read it, distrust it.
- 12The write path: one row per span, one key per traceSpans are written as they arrive under a random trace id that spreads perfectly across partitions. Nothing ever closes a trace, so a partial trace is the normal steady state and the tree is rebuilt at read time.~7 min
- 13Nobody knows the trace idEvery way into the store other than the id costs something: an index prices high-cardinality attributes, a scan trades that storage for query compute. The cheapest way in is to arrive already holding the id.~7 min
- 14Where did the latency go? — critical path and self time in a waterfallThe waterfall answers three questions at a glance: which span holds time its children don't account for, which children overlap, and where there is no span at all. Only the chain that doesn't overlap can be shortened.~7 min
- 14aWhen the picture itself lies — clock skew, gaps and the uninstrumented hopEvery span's times come from the clock of the machine that recorded it, and those clocks disagree. Viewers quietly shift timestamps to make a child fit its parent, folding real network and queue wait into the parent.~7 min
Ship it
What a million traces say, then design the stack.
- 15What a million traces say together — the service dependency graph and span metricsThe parent-child edges of many traces draw the live dependency graph, and counting spans by service and operation yields rates and error ratios for hops no dashboard covered — both only as honest as the sampling.~7 min
- 16Design your tracing stackA tracing stack is a chain of decisions: what to instrument, what to propagate, how much to keep and by what rule, where that decision is made, and how a trace gets found. Each is right only for a given workload.~7 min
More in Observability: Metrics, Logs & Traces
The three pillars, built rather than bought: a time-series database, a log pipeline, a tracing system, and the metrics platform on top.
- Build Metrics / Monitoring SystemTime-series at scale. Cardinality is the enemy.
- Build Build a distributed logging stack (ELK / Loki)Ship lines off N hosts, choose what to index, age data through tiers, retain or delete on schedule, and survive a chatty service — built one decision at a time.
- Build Build a Prometheus-style time-series databaseThe simplest database that can absorb a 1M-points-per-second firehose and still answer `sum(rate(http_requests_total{status="500"}[5m]))` in milliseconds — built bit by bit, literally.
Prefer to design it yourself?
The same subject as a staged workspace: draw the architecture, and a simulator traces requests through the boxes you drew.
Open the Build a distributed tracing system (Jaeger / Zipkin style) workspace