Single-node by design; HA is somebody else's problem

The TSDB itself is explicitly single-node — no consensus, no replication, no leader election; HA is two identical scrapers running side-by-side, and durable cluster storage is delegated to a separate remote-write tier that dedups by __replica__ label.

Previously

The single-node TSDB is a complete unit. Scaling out is a separate tier with separate trade-offs.

Scene 11

Single-node by design; HA is somebody else's problem

  1. Watch
  2. Try it
  3. Predict
  4. Capture
1000 scrape targetsevery 15sthe source of truthSCRAPE STREAMscrape-targetstopictwo replicas read independently — no replication in TSDBHA = DUPLICATE SCRAPERS + REMOTE-WRITE# Two HA replicas scrape the SAME targetspromA.scrape(targets, replica='A')promB.scrape(targets, replica='B')# Each ships every sample via remote_writeremote_write(sample, __replica__='A')remote_write(sample, __replica__='B')# Receiver dedups: keep one, drop the otherif seen(sample.fingerprint): drop()Prometheus A__replica__=A · offset 0Prometheus B__replica__=B · offset 0Remote-write receiverdedup by __replica__ · 0 duplicates dropped
What to watch for

Two Prometheus instances scrape the same 1000 targets — that's the entire HA story. Both ship every sample to the remote-write receiver downstream. The TSDB itself doesn't replicate, so how do you survive a node dying? Run two of them, scraping the same targets, and ship every sample to a separate cluster. That ship is called remote write. The receiver sees each sample twice (with __replica__=A and __replica__=B) and drops the duplicate.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Prometheus.run
every interval: scrape, append locally, queue remote-write
def run(self):
while alive:
sleep(scrape_interval) # 15s default
for target in self.targets: # SAME targets as peer
samples = target.scrape()
for s in samples:
s.labels['__replica__'] = self.replica_id
self.tsdb.append(s) # local single-node store
self.remote_write_queue.put(s)
RemoteWriteReceiver.handle
key = (series, ts); if seen with different __replica__, drop
def handle(self, sample):
fp = fingerprint(
sample.metric, sample.labels_without_replica,
)
key = (fp, sample.timestamp)
prev = self.seen.get(key)
if prev is None:
self.seen[key] = sample.labels['__replica__']
self.store.append(sample) # durable cluster store
return
if prev != sample.labels['__replica__']:
self.dedup_drops += 1 # duplicate from peer

Where this sits in Build a Prometheus-style time-series database

Scene 10a of 12. The TSDB itself isn't replicated. HA = two parallel scrapers; durability = ship every sample to a remote-write receiver that dedups.

Up next. We have all the pieces. Now you build: pick a workload, pick a configuration, and watch the simulator tell you whether it survives.

All 12 scenes in Build a Prometheus-style time-series database · Every curriculum

Built with Arqly
Every scene in Build a Prometheus-style time-series database builds on the one before it.All 12 Build a Prometheus-style time-series database scenes