Sampling — head, tail, and the only error — per-tenant ingest quotas and 429 admission control

Head sampling decides at emit time (cheap, can't condition on outcomes); tail sampling decides at the collector after buffering the full trace (can keep all errors, but must hold every in-flight record and route consistently) — and a uniform 1% sample over a mixed stream will, with overwhelming probability, drop the one error you needed.

Previously

Bytes die on a schedule and the index has to forget. But even with perfect retention, at scale we cannot keep everything we emit. The honest question is which lines we sacrifice — and the wrong answer is 'the only error of the day'.

Scene 09

Sampling — head, tail, and the only error

  1. Watch
  2. Try it
  3. Predict
  4. Capture
uniform head · p=0.01INCOMING STREAM · 10,000 lines/min · errorRate = 0.000%infodbginfoinfoinfodbgWRNinfoinfodbginfoinfoinfodbginfoinfoinfoERRinfoinfoinfodbginfo↑ the only ERRORHEAD SAMPLERfilter p=0.01 (uniform)p=0.01dropped at agent→ backendTAIL SAMPLERdecision_wait = 30s · buffer = 60 MBcollector buffer60 MBkeep if any record.level == E…→ backenderrors observed @ backend (head)1.0%errors observed @ backend (tail)0/24ERROR-CAPTURE RATEP(≥1 captured in 1h, head) = 1.0%head1.0%tail100.0%RATE LIMIT LANE(separate from sampling)Loki ingestion_rate_mb /ELK bulk-queuequota4.0 MB/s3.00 MB/s (75%)within quotasampling ≠ rate limitingHead sampling at p=0.01 over a mixed INFO+ERROR stream. With 1 error/hour, P(captured per error) = 1%. Over a…
incoming stream — one rare red ERROR mixed in
What to watch for

Watch the stream at the top. Most records are INFO; one rare red ERROR is mixed in around the middle. The stream forks LEFT into a HEAD SAMPLER (uniform p=0.01 at the agent) and RIGHT into a TAIL SAMPLER (collector buffers for 30s then keeps the trace if any record was an ERROR). The capture-rate meter at the bottom tells the story: head drops the rare red cell; tail keeps it. To the right, a separate lane shows the rate-limit mechanism — per-tenant quota with HTTP 429s — distinct from sampling, same diagram.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Agent.head_sample
decision at emit time on the agent — no buffer, blind to outcomes
# uniform: every record gets the same coin flip
def head_sample_uniform(line, p=0.01):
if random() < p:
ship(line)
else:
drop(line) # never reaches the backend
# level-aware: per-level p map (needs structured logs)
P_BY_LEVEL = {info: 0.01, debug: 0.01,
warn: 1.0, error: 1.0}
def head_sample_level_aware(line):
p = P_BY_LEVEL[line.level]
if random() < p: ship(line) else: drop(line)
Collector.tail_sample
decision at the collector after buffering decision_wait
buffer = {} # trace_id -> list[span] (per collector)
def on_span(span):
# consistent routing required: all spans of a trace
# MUST land at the same collector or the rule sees
# a partial trace.
buffer.setdefault(span.trace_id, []).append(span)
def on_idle(trace_id, decision_wait=30):
spans = buffer.pop(trace_id)
if any(s.level == ERROR for s in spans):
ship(spans) # keep on error
else: # drop the whole trace
drop(spans)
Backend.rate_limit
per-tenant ingestion_rate_mb — admission control, not sampling
# Loki: ingestion_rate_mb / ingestion_burst_size_mb
# ELK: bulk thread pool fills -> bulk-queue-rejected
def on_push(tenant, batch):
used = tokens_used(tenant) # bytes/sec window
if used + batch.bytes > ingestion_rate_mb:
return HTTP_429 # agent retries / spills to disk
accept(batch)
advance_tokens(tenant, batch.bytes)

Where this sits in Build a distributed logging stack (ELK / Loki)

Scene 09 of 12. Head sampling decides at emit (cheap, blind); tail sampling decides at the collector (can keep all errors, costs buffer). A uniform 1% sample drops the only error you needed.

Up next. We've shaped the data on its way in (sampling), through (indexing), and out (retention). The backend that runs all this is itself a distributed system — briefly, the moving parts.

All 12 scenes in Build a distributed logging stack (ELK / Loki) · Every curriculum

Built with Arqly
Every scene in Build a distributed logging stack (ELK / Loki) builds on the one before it.All 12 Build a distributed logging stack (ELK / Loki) scenes