Nobody knows the trace id
Every way into the store other than the id costs something — a second structure over span fields prices high-cardinality attributes the way a log index does, a scan trades that storage for query compute — so the cheapest way in is to arrive already holding the id.
Get-by-id is one cheap lookup. The engineer looking at a page does not have an id — they have a symptom and a time range.
Scene 13
Nobody knows the trace id
- Watch
- Try it
- Predict
- Capture
It is 3 a.m., the checkout alert just fired, and you have a symptom and a time range — not a trace id. So how do you find the one trace that matters? Watch three routes go after the same 3-second Shopfront checkout. The first searches a second structure the store keeps over span fields — service, operation, status, duration — which is called a span index. The second keeps no such structure and reads the raw blocks in the window. The third does not search at all.
Highlighted lines are the ones running in the diagram right now.
# a budget decision, not a default:INDEXED_FIELDS = intrinsics + nominated_attrs# intrinsics = service, operation, status, durationdef write_span(span):blocks.append(span)for field in INDEXED_FIELDS:value = span.get(field)if value is None: continue# one posting per distinct value of fieldindex.add(field, value, span.trace_id)# index compacts and is backed up with the blocks
def by_attribute(question, window):if not question.fields <= INDEXED_FIELDS:return CANNOT_ANSWERpostings = index.lookup(question.fields, question.values,)# window narrows postings; it reads nothing moreids = [p.trace_id for p in postingsif p.time in window]return [get_by_id(tid) for tid in ids]
def scan_blocks(question, window):candidates = []for block in blocklist:# every block records its min/max timeif block.max_time < window.start: continueif block.min_time > window.end: continueif not block.stats.may_contain(question):continuecandidates.append(block)# thousands of short-lived jobs, a slice eachreturn fan_out(candidates, read_and_match)
# wired up on the metrics side, months agodef observe(duration, ctx):bucket = histogram.bucket_for(duration)bucket.count += 1 # the value counts here toobucket.exemplars.append(Exemplar(trace_id=ctx.trace_id,span_id=ctx.span_id,value=duration,)) # up to 5 per data pointdef open_from_p99(bucket): # no search runsreturn get_by_id(bucket.exemplars[0].trace_id)
Where this sits in Build a distributed tracing system (Jaeger / Zipkin style)
Scene 13 of 17, in the Store & read act — Key by trace id, find it, read it, distrust it.. Every way into the store other than the id costs something: an index prices high-cardinality attributes, a scan trades that storage for query compute. The cheapest way in is to arrive already holding the id.
Up next. You are now holding the right trace. What do you actually look at to name the culprit in under five minutes?
All 17 scenes in Build a distributed tracing system (Jaeger / Zipkin style) · Every curriculum