Sampling that doesn't lie — sampling bias and adjusted counts

A uniform sample keeps proportions roughly right but destroys rare events, a per-service trace budget quietly over-represents the quiet services, and counting stored traces without multiplying each by one over its keep rate understates everything by exactly the sampling factor.

Previously

Consistent sampling gives you whole traces. It does not give you a representative answer to every question you might ask of them.

Scene 08

Sampling that doesn't lie

  1. Watch
  2. Try it
  3. Predict
  4. Capture
what the kept sample can still honestly saykeep the same fraction of everythingraw sample countscheckout · ok, under 2 ssame scale, both barstruth9380sample0checkout · ok, over 2 ssame scale, both barstruth220sample0checkout · failedsame scale, both barstruth400sample0admin · all requestssame scale, both barstruth45sample0raw sample counts — a service that kept less simply looks smallercounted from the kept traces—counted from metrics—vstruth—the one checkout in ten thousand that hit the 9-second fraud-check timeout? never recordedhow many is a question for metrics (metrics-05); a trace explains one of the requests you happened to keepTIME TO CULPRIT—One hour of Shopfront checkout traffic (example). Each row's upper bar is what really happened.
What to watch for

You kept 1 trace in 100. Which questions can those traces still answer honestly, and which ones can they no longer answer at all? This is one hour of Shopfront checkout traffic (example). Watch the upper bar of each pair — what really happened — and then the lower bar that a uniform 1-in-100 sampler kept, drawn on the same scale. Two of the rows shrink neatly in proportion. Watch what happens to the other two. Then watch the failure count get taken twice, once from the kept traces and once from metrics, and notice that nothing on the diagram explains why they disagree.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Sampler.should_sample
three ways to choose, three different p
def should_sample(request):
if strategy == "uniform":
return keep(request, p = 0.01)
if strategy == "rate-limit-per-service":
bucket = buckets[request.service]
# N traces/sec, so throughput sets p
return keep(request, p = bucket.rate())
if strategy == "keep-errors-plus-baseline":
# needs the outcome, not known here yet
if request.failed:
return keep(request, p = 1.0)
return keep(request, p = 0.01)
Sampler.keep
the rate rides along with the trace it decided
def keep(request, p):
if not coin_flip(p):
return DROP
trace = start_trace(request)
# the rate this one trace survived at
trace.keep_rate = p
store.write(trace)
return RECORD
Counter.estimate
counting stored traces vs counting what they stand for
def estimate(query):
stored = store.select(query)
if not weight_by_keep_rate:
return len(stored)
total = 0
for trace in stored:
total += 1 / trace.keep_rate
return total
def estimate_from_one_rate(stored):
return len(stored) / assumed_rate

Where this sits in Build a distributed tracing system (Jaeger / Zipkin style)

Scene 08 of 17, in the Sampling act — You can't keep it all; decide once, and don't lie.. A uniform sample keeps proportions but destroys rare events; rate limiting over-represents quiet services; and counting stored traces without weighting each by the inverse of its keep rate understates everything.

Up next. Reweighting fixes the counts, but nothing can resurrect the one failure that was never kept. What if the decision waited until the trace was over?

All 17 scenes in Build a distributed tracing system (Jaeger / Zipkin style) · Every curriculum

Built with Arqly
Every scene in Build a distributed tracing system (Jaeger / Zipkin style) builds on the one before it.All 17 Build a distributed tracing system (Jaeger / Zipkin style) scenes