Every request, every hop, forever? — full capture cost and the keep rate

Spans multiply — requests per second times hops per request — so full capture is a storage bill measured in rows per day plus a measured latency cost at the source, which makes traces the one pillar where you deliberately keep a fraction.

Previously

An attribute costs bytes on a row that already exists. The question that raises is how many rows there are.

Scene 06

Every request, every hop, forever?

  1. Watch
  2. Try it
  3. Predict
  4. Capture
what a kept trace costs: hops multiply, then the keep rate dividesrequests into Shopfront2k/sexamplespans recorded—raw, before any choice—written to the backend—keep rate — what the sampler lets through100%0% of traces are never written anywherestored per day (GB) · 500 B a span0budgetof yesterday's 40 failed checkouts, this many still have a trace you can open— / —kept · out ofTIME TO CULPRIT—One request does not cost one row. It costs one row per hop, every second, every day.
What to watch for

How much does it actually cost to record every request, at every hop, and keep it? Shopfront takes about 2,000 checkout-path requests a second (example), and each one crosses 8 services — so each one is 8 spans, not 1. Watch the chain multiply: requests per second, then hops, then bytes per stored span. The last card is what lands on disk every day, and the dashed line is the team's daily allowance. Keep an eye on the amber strip at the bottom too: the second cost of recording everything is not paid on disk at all.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Capacity.raw_per_day
what recording every hop costs before any choice
BYTES_PER_SPAN = 500 # Tempo plans ~500 B; Dapper: 426 B
SECONDS_PER_DAY = 86_400
def raw_per_day(requests_per_sec, hops_per_request):
spans_per_sec = requests_per_sec
spans_per_sec *= hops_per_request # one span per hop
bytes_per_sec = spans_per_sec * BYTES_PER_SPAN
return bytes_per_sec * SECONDS_PER_DAY
raw = raw_per_day(rate_into_checkout, services_on_path)
Capacity.stored_per_day
the one fraction that divides the bill, and what else it divides
DAILY_BUDGET = 100 * GB
def stored_per_day(raw, keep):
# sampling: record only this fraction of traces
return raw * keep
def over_budget(raw, keep):
return stored_per_day(raw, keep) > DAILY_BUDGET
def still_have_a_trace(rare_events, keep):
return rare_events * keep
BatchProcessor.on_end
the second line of the bill, paid in latency
def on_end(span): # still on the request path
buf = serialize(span) # the term that costs
queue.put(buf) # a background thread batches
# the span object itself: 204 ns root, 176 ns non-root
# Dapper's web-search cluster, measured per rate:
# 1/1 +16.3% latency, -1.48% throughput
# 1/8 +4.12%, -0.23%
# 1/16 +2.12%, -0.08%

Where this sits in Build a distributed tracing system (Jaeger / Zipkin style)

Scene 06 of 17, in the Sampling act — You can't keep it all; decide once, and don't lie.. Spans multiply: requests per second times hops per request. Full capture is a storage bill plus a measured latency cost on the request path, which makes traces the one pillar where you deliberately keep a fraction.

Up next. You cannot keep everything, so you have to keep a fraction. Who decides which fraction — and what happens if every service decides for itself?

All 17 scenes in Build a distributed tracing system (Jaeger / Zipkin style) · Every curriculum

Built with Arqly
Every scene in Build a distributed tracing system (Jaeger / Zipkin style) builds on the one before it.All 17 Build a distributed tracing system (Jaeger / Zipkin style) scenes