Build your own distributed logging stack (ELK / Loki)
Ship lines off N hosts, choose what to index, age data through tiers, retain or delete on schedule, and survive a chatty service — built one decision at a time.
- Scenes
- 12 interactive scenes
- Time
- about 84 minutes
- Topic
- Observability: Metrics, Logs & Traces
What you are building, and why
You are building the simplest stack that can take log lines emitted by hundreds of pods on dozens of hosts, get them off-host before they age out of the on-host file, store them somewhere a query can reach, and answer "show me every ERROR matching 'connection refused' in the last hour" in a time the on-call engineer is willing to wait. Clients emit lines (log.info(...)); the system tails, batches, ships, parses, indexes, tiers, retains, and queries them; somewhere in the middle a thoughtful operator decides which lines are worth indexing and which can stay opaque.
ELK and Loki are the canonical worked examples — opposite ends of the same trade-off curve. Once you build the core, the design space of every logging stack collapses to a handful of named trades — pay-at-write vs pay-at-read, the cardinality ceiling you defend, the cost ladder you age through, and where you put the line between "we keep it" and "we sampled it away."
Resist the urge to "describe ELK" or "describe Loki." Make decisions yourself, defend them, and let the workload push back.
What you will be able to explain afterwards
- the shipping agent: tail offset, batch, retry, backpressure
- block / drop / spill: how a backend outage starves application threads
- structured logs: fields (ELK) vs labels (Loki), parse-at-write vs parse-at-read
- the indexing fork: Lucene inverted index (heavy index, fast read) vs Loki labels-only index + chunks (cheap write, grep-on-read)
- execution-plan asymmetry: posting-list intersection vs label-resolve-then-stream-from-S3
- cardinality: low-card → labels/fields, high-card → body / structured metadata
- tiering: hot/warm/cold/frozen ILM and the two-orders-of-magnitude cost ladder
- retention vs deletion: tombstones, force-merge lag, the PII per-template leak
- head sampling vs tail sampling: when uniform sampling drops the only error
- backend distribution: distributors, ingesters, queriers — partition by hash(doc_id) vs hash(label-set)
- 01grep + ssh stops working — centralized log collection across a host fleetOnce your fleet has more than a handful of hosts, the only way to answer 'which host saw this error' is to ship lines centrally — local files plus ssh is an O(N) dead end.~7 min
- 02An agent on every host — at-least-once log shipping and offset checkpointsA shipping agent tails each log file from a saved offset, batches lines, and POSTs them at-least-once — duplicates on retry are normal, not a bug.~7 min
- 03Block, drop, or spill — when_full=block paired with a synchronous loggerWhen the backend stalls, the agent must block, drop, or spill to disk — and `when_full=block` plus a synchronous logger is how a logging outage takes the application down with it.~7 min
- 04String vs map — fields and labels — structured logging and write-time JSON parsingA log line is either a string parsed at read-time or a typed map parsed at write-time, and the two systems we'll meet attach different names to the same idea — fields in ELK, labels in Loki.~7 min
- 05Inverted index vs labels-only indexELK tokenises every value into a per-term posting list; Loki hashes the label-set into a stream id and appends the line verbatim — heavy index + small body vs tiny index + verbatim body.~7 min
- 05aSame query, two execution plansELK answers via posting-list intersection — milliseconds; Loki resolves labels to chunks, fetches them from S3, and greps in-process — seconds to minutes. Opposite ends of the same trade-off curve.~7 min
- 06Cardinality is the killerEvery unique label-set is a Loki stream; every dynamic key is an ELK mapping field. Putting request_id in either kills the index in minutes — low-card → labels, high-card → body.~7 min
- 07Hot, warm, cold, frozen — log storage tiers and ILM phasesTwo orders of magnitude in cost between NVMe and Deep Archive force tiering. ELK has four ILM phases; Loki collapses to S3 from day one.~7 min
- 08Retention vs deletion — the index has to forget — delete-by-query tombstones and manual force-mergeRetention is when the system stops promising you can read; deletion is when bytes are physically gone — and the gap is where compliance bugs live.~7 min
- 09Sampling — head, tail, and the only error — per-tenant ingest quotas and 429 admission controlHead sampling decides at emit (cheap, blind); tail sampling decides at the collector (can keep all errors, costs buffer). A uniform 1% sample drops the only error you needed.~7 min
- 10Distributors, ingesters, queriers (briefly)Both ELK and Loki are sharded write-paths plus sharded read-paths plus a backing store; they differ in the partition key — hash(doc_id) vs hash(label-set).~7 min
- 11Design your logging stackCapstone: agent + buffer policy + structuring + index strategy + tiers + retention + sampling — the verifier traces every choice back to the scene that earned it.~7 min
More in Observability: Metrics, Logs & Traces
The three pillars, built rather than bought: a time-series database, a log pipeline, a tracing system, and the metrics platform on top.
- Build Metrics / Monitoring SystemTime-series at scale. Cardinality is the enemy.
- Build Build a distributed tracing system (Jaeger / Zipkin style)Metrics tell you that checkout is slow and logs tell you what one service printed; neither can say which hop of one request consumed the time. Build the third pillar: one record per hop tied together by one id — span propagation, sampling that doesn't lie, an export path that never blocks a customer, storage keyed by trace id, and the waterfall that finally answers 'where did the latency go?'
- Build Build a Prometheus-style time-series databaseThe simplest database that can absorb a 1M-points-per-second firehose and still answer `sum(rate(http_requests_total{status="500"}[5m]))` in milliseconds — built bit by bit, literally.
Prefer to design it yourself?
The same subject as a staged workspace: draw the architecture, and a simulator traces requests through the boxes you drew.
Open the Build a distributed logging stack (ELK / Loki) workspace