High cardinality is free here

A span is already one row per request per hop, so attaching the customer id adds bytes to a row that was being written anyway and multiplies nothing — the exact opposite of a metric label, where each distinct value creates another series that lives for the whole retention window.

Previously

Every hop now produces a span in the right trace, including across the queue. A span is a row that is being written anyway — so what else can ride on that row?

Scene 05

High cardinality is free here

  1. Watch
  2. Try it
  3. Predict
  4. Capture
one span, filling up — and the same value on the other sidePOST /payments/chargepaymentsONE SPAN — ONE ROWATTRIBUTES0 fieldsEVENTSnone recorded on this span186 B on this row — Dapper measured 426 B per stored span186 B / 5.1 KBthe row was being written anyway — attributes only add bytes to itTHE SAME VALUE, TWO SIDEScom.shopfront.customer_idnot added yetAS A SPAN ATTRIBUTE1rows, whatever the value isAS A METRIC LABEL1series, one per distinct valueWhy each distinct value costs memory for the whole retention window: metrics-02.WHAT DOES COST HEREbytes on every span, every hop426 B → 1 KB+credentials written down by accident1 fieldan index built over this key laterpay then, not nowA metric label multiplies storage by the number of distinct values; a span attribute multiplies it by the number of bytes.
What to watch for

A span is already being written for this hop — so what else can ride on it, and what does riding on it cost? Watch one payments span fill with rows. Each row is a key and a value; the meter at the foot of the card climbs as they land. Keep one eye on the counter to the right of the card: it stays at 1 the whole time, because this is one row, not a series. At the end, one more row lands underneath the others, with a timestamp on it.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Span.set_attribute
one more key/value on a row that already exists
def set_attribute(self, key, value):
# keys: lowercase, dot-delimited; your own
# take a com.* prefix — otel.* is reserved
if is_credential(key, value):
value = redact(value) # url.full
# attr_count_limit defaults to 128
if len(self.attributes) >= attr_count_limit:
self.dropped_attributes_count += 1
return # silent, but counted
self.attributes[key] = value
self.bytes += len(key) + len(value)
Span.add_event
a timestamped moment inside this hop
def add_event(self, name, attributes, at):
if len(self.events) >= event_count_limit:
self.dropped_events_count += 1
return
self.events.append(Event(name, at, attributes))
# no new span_id, no new parent_span_id —
# the trace keeps the shape it already had
def record_exception(self, err):
self.add_event("exception",
exception_attrs(err), clock.now())
Handler.on_charge
the customer id written to the span and to a metric
def on_charge(span, customer_id):
# the same value, on both sides at once
span.set_attribute(
"com.shopfront.customer_id", customer_id)
charges.add(1, {"customer_id": customer_id})
# what charges.add does in the metrics SDK
def add(self, n, labels):
key = fingerprint(labels)
if key not in self.series:
self.series[key] = new_series(key)
self.series[key].value += n

Where this sits in Build a distributed tracing system (Jaeger / Zipkin style)

Scene 05 of 17, in the Carry it act — The header, the queue that drops it, what rides along.. A span is already one row per request per hop, so attaching a customer id adds bytes to a row that was being written anyway and multiplies nothing — the exact opposite of a metric label.

Up next. Attributes are nearly free per span, but a span itself is not free, and there is one per hop per request. What does keeping every one of them actually cost?

All 17 scenes in Build a distributed tracing system (Jaeger / Zipkin style) · Every curriculum

Built with Arqly
Every scene in Build a distributed tracing system (Jaeger / Zipkin style) builds on the one before it.All 17 Build a distributed tracing system (Jaeger / Zipkin style) scenes