Design your tracing stack
A tracing stack is a chain of decisions — what to instrument, what to propagate, how much to keep and by what rule, where that decision is made, how a trace gets found, and what may be derived from it — and each is right only for a given workload.
Every mechanism in the curriculum now has a price attached. The only remaining question is which ones a given workload actually needs.
Scene 16
Design your tracing stack
- Watch
- Try it
- Predict
- Capture
Which of the fifteen mechanisms you have met does a given workload actually need? Here is a small one: three services, 200 requests a second, traces read by hand while debugging. A stack is already placed, and the checker on the right is green — but look at what makes each verdict trustworthy: every line names the scene it came from. Then watch a gateway tail-sampling tier get added to the design, and watch the checker's answer change without a single fact about the workload changing.
Highlighted lines are the ones running in the diagram right now.
MAX_TRACE_RETENTION = 30 * DAY # X-Ray, fixeddef check(design, workload):if workload.retention > MAX_TRACE_RETENTION:return REFUSE(asked_for = 'every payment, kept forever',belongs_in = 'append-only record, by account',)return [judge_chain(design, workload), # 03, 04judge_keep(design, workload), # 06, 07, 09, 11judge_entry(design, workload), # 10, 13, 15]
def judge_chain(design, workload):for hop in workload.hops:if hop.kind == 'http':carried = design.header('traceparent') # 03else:# no outbound request to put a header oncarried = design.inject_into_message() # 04if not carried:return fail('the chain stops at this hop;'' the next service starts a trace')return ok()
def judge_keep(design, workload):if workload.affordable_at_full_keep: # 06if design.sampling == 'tail':return warn('nothing for a policy to choose')return ok('keep 100%; deciding costs more')out = []if workload.must_keep('errors', 'over 2 s'): # 09if design.sampling != 'tail':out.append(fail('decided at the first span'))if design.decide_at != 'gateway_pool': # 07, 11out.append(fail('one decision, on a whole trace'))return out or [ok()]
SPAN_BYTES = 500 # a span on the wirePER_MILLION_TRACES = 5.00 # X-Ray, recordeddef sizing(design, workload):spans = workload.rps * workload.hops_per_requestingest = spans * SPAN_BYTES # tail ships all of itstored_per_day = ingest * design.keep_rate * 86400bill = workload.rps * 86400 / 1e6 * PER_MILLION_TRACEStiers = 2 if design.decide_at == 'gateway_pool' else 1held_for = min(workload.retention, MAX_TRACE_RETENTION)return stored_per_day, bill, tiers, held_for
Where this sits in Build a distributed tracing system (Jaeger / Zipkin style)
Scene 16 of 17, in the Ship it act — What a million traces say, then design the stack.. A tracing stack is a chain of decisions: what to instrument, what to propagate, how much to keep and by what rule, where that decision is made, and how a trace gets found. Each is right only for a given workload.
All 17 scenes in Build a distributed tracing system (Jaeger / Zipkin style) · Every curriculum