Decide once, carry the answer — head-based sampling and the sampled flag
Sampling is per trace, not per span: if N hops each flip their own coin at probability p, a complete trace survives with probability p to the power N, and what you store is mostly fragments that look like whole traces.
Keeping a fraction is settled. Who makes that call, and where, is not — and the obvious answer is that each service decides what it can afford.
Scene 07
Decide once, carry the answer
- Watch
- Try it
- Predict
- Capture
If you are only going to keep some traces, who makes that call — and where? The obvious answer is the one every team reaches for first: let each service keep the fraction it can afford. So all five hops of one Shopfront checkout are set to 10%, and each one flips its own coin the moment the request lands on it. Watch the lanes thin, one hop at a time, and watch the two counters at the foot: how many checkouts end up recorded at every single hop, and how many end up recorded at some hops but not others. You already met keeping-a-fraction when we sampled log lines in the logging curriculum, and none of that is re-derived here. Two things are new: the unit is a whole trace rather than a single line, and every hop has to land on the same answer as every other hop.
Highlighted lines are the ones running in the diagram right now.
def should_sample(parent_ctx, trace_id):# parent_ctx is never read, so a hop cannot# see what any other hop decidedreturn random() < self.ratiobase = self.ratio # configured per serviceexponent = len(path) # one call per hopwhole = base ** exponent
def should_sample(parent_ctx, trace_id):if parent_ctx is None: # only hop 1 lands herereturn self.root.should_sample(trace_id)if parent_ctx.sampled: # the bit off traceparentreturn True # record, unconditionallyreturn False# self.ratio is an argument to self.root, so a hop# that already has a parent never consults it# one decision per trace, whatever len(path) is# OTEL_TRACES_SAMPLER=parentbased_traceidratio
def should_sample(parent_ctx, trace_id):# every hop already carries the trace idr = randomness_from(trace_id) # 0.0 to 1.0return r < self.ratio# no coin, no flag, nothing negotiated: the same id# and the same ratio give the same answer at every# hop on the path# so every service has to be given the same ratio# OTEL_TRACES_SAMPLER=traceidratio
Where this sits in Build a distributed tracing system (Jaeger / Zipkin style)
Scene 07 of 17, in the Sampling act — You can't keep it all; decide once, and don't lie.. Sampling is per trace, not per span. If N hops each flip their own p coin a whole trace survives with probability p to the N, and what you store is mostly fragments that look like whole traces.
Up next. One decision per trace gives you whole traces. Now that the traces you kept are real, what can you honestly say about the ones you did not keep?
All 17 scenes in Build a distributed tracing system (Jaeger / Zipkin style) · Every curriculum