Build a distributed tracing system (Jaeger / Zipkin style)
Metrics tell you that checkout is slow and logs tell you what one service printed; neither can say which hop of one request consumed the time. Build the third pillar: one record per hop tied together by one id — span propagation, sampling that doesn't lie, an export path that never blocks a customer, storage keyed by trace id, and the waterfall that finally answers 'where did the latency go?'Enter to send · Shift+Enter for a new line
About Build a distributed tracing system (Jaeger / Zipkin style)
Metrics tell you that checkout is slow and logs tell you what one service printed; neither can say which hop of one request consumed the time. Build the third pillar: one record per hop tied together by one id — span propagation, sampling that doesn't lie, an export path that never blocks a customer, storage keyed by trace id, and the waterfall that finally answers 'where did the latency go?'
- Difficulty
- intermediate
- Time
- about 102 minutes
- Stages
- 9
- Topic
- Observability: Metrics, Logs & Traces
How this problem is worked
Nine stages, from what the thing is for to how it compares with the real implementations. Each asks one question, and the simulator runs the architecture you draw against the requirements you wrote.
- 01Purpose & invariantsWhat is this for, and what must always be true of it?
- 02Workload characterizationWho writes, who reads, and in what shapes?
- 03Data model & on-disk formatWhat does the data look like at rest?
- 04Core algorithmsHow do the write path and the read path actually work?
- 05Distribution & replicationHow does this scale out and survive losing a machine?
- 06Consistency & correctnessUnder concurrency and failure, what is guaranteed?
- 07Failure modes & recoveryWhat actually happens when each part fails?
- 08Operational characteristicsCan a human run this at three in the morning?
- 09Trade-offs & comparisonWhere does this sit against the alternatives?
Primary sources for this problem
- Sigelman et al. — Dapper, a Large-Scale Distributed Systems Tracing Infrastructure (Google 2010)
- W3C Trace Context Recommendation — traceparent, tracestate, and the flags byte
- OpenTelemetry specification — trace API/SDK, samplers, propagators, OTLP
- OpenTelemetry Collector — tail_sampling processor, loadbalancing exporter, spanmetrics connector
- Jaeger documentation — architecture, sampling, storage backends, clock-skew adjustment
- Grafana Tempo documentation — Parquet block format, TraceQL, backend search
- Dean & Barroso — The Tail at Scale (CACM 2013)
- Uber Engineering — CRISP: critical-path analysis for microservice architectures
More in Observability: Metrics, Logs & Traces
The three pillars, built rather than bought: a time-series database, a log pipeline, a tracing system, and the metrics platform on top.
- Build Metrics / Monitoring SystemTime-series at scale. Cardinality is the enemy.
- Build Build a distributed logging stack (ELK / Loki)Ship lines off N hosts, choose what to index, age data through tiers, retain or delete on schedule, and survive a chatty service — built one decision at a time.
- Build Build a Prometheus-style time-series databaseThe simplest database that can absorb a 1M-points-per-second firehose and still answer `sum(rate(http_requests_total{status="500"}[5m]))` in milliseconds — built bit by bit, literally.
Browse the full problem catalog, or see what the simulator does and does not model.