Sampling that doesn't lie — sampling bias and adjusted counts
A uniform sample keeps proportions roughly right but destroys rare events, a per-service trace budget quietly over-represents the quiet services, and counting stored traces without multiplying each by one over its keep rate understates everything by exactly the sampling factor.
Consistent sampling gives you whole traces. It does not give you a representative answer to every question you might ask of them.
Scene 08
Sampling that doesn't lie
- Watch
- Try it
- Predict
- Capture
You kept 1 trace in 100. Which questions can those traces still answer honestly, and which ones can they no longer answer at all? This is one hour of Shopfront checkout traffic (example). Watch the upper bar of each pair — what really happened — and then the lower bar that a uniform 1-in-100 sampler kept, drawn on the same scale. Two of the rows shrink neatly in proportion. Watch what happens to the other two. Then watch the failure count get taken twice, once from the kept traces and once from metrics, and notice that nothing on the diagram explains why they disagree.
Highlighted lines are the ones running in the diagram right now.
def should_sample(request):if strategy == "uniform":return keep(request, p = 0.01)if strategy == "rate-limit-per-service":bucket = buckets[request.service]# N traces/sec, so throughput sets preturn keep(request, p = bucket.rate())if strategy == "keep-errors-plus-baseline":# needs the outcome, not known here yetif request.failed:return keep(request, p = 1.0)return keep(request, p = 0.01)
def keep(request, p):if not coin_flip(p):return DROPtrace = start_trace(request)# the rate this one trace survived attrace.keep_rate = pstore.write(trace)return RECORD
def estimate(query):stored = store.select(query)if not weight_by_keep_rate:return len(stored)total = 0for trace in stored:total += 1 / trace.keep_ratereturn totaldef estimate_from_one_rate(stored):return len(stored) / assumed_rate
Where this sits in Build a distributed tracing system (Jaeger / Zipkin style)
Scene 08 of 17, in the Sampling act — You can't keep it all; decide once, and don't lie.. A uniform sample keeps proportions but destroys rare events; rate limiting over-represents quiet services; and counting stored traces without weighting each by the inverse of its keep rate understates everything.
Up next. Reweighting fixes the counts, but nothing can resurrect the one failure that was never kept. What if the decision waited until the trace was over?
All 17 scenes in Build a distributed tracing system (Jaeger / Zipkin style) · Every curriculum