Pull or push: who starts the conversation — scrape targets, the up metric and service discovery

When the monitor pulls each target on its own schedule, a failed fetch itself says the target is down, while when apps push, short-lived jobs can still report but silence becomes ambiguous: dead, or just quiet?

Previously

Totals survive missed reads; now we decide who starts each read and what a failed read tells us.

Scene 04

Pull or push: who starts the conversation

  1. Watch
  2. Try it
  3. Predict
  4. Capture
measurecollectmonitor fetcheskeepquerychecknotifyTIME TO DETECT≈ 30 sexampleKubernetes pod listwho exists right nowlist of podsmonitorfetches every 15 s (example)pod-1/metricspod-2/metricspod-3/metricspod-4/metricspod-5/metricspod-6/metricspod-7/metricspod-8/metricspod-9/metricspod-10/metricsbatch job: 3 snever seen0 s60 seach tick = one fetch, every 15 sThe monitor asks Kubernetes which pods exist, then fetches each one's page of numbers.
What to watch for

How do the numbers get from the app to the monitor, and what does each direction cost? Here the monitor does the asking. It first gets Kubernetes' list of checkout pods that exist right now (finding targets this way is called service discovery), then walks that list, fetching each pod's page of numbers every 15 s (example). Watch the second round: pod-7 hangs while Kubernetes still lists it.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Monitor.scrape
the monitor asks every listed target for its page
def scrape_all(): # every 15 s (example)
targets = kubernetes.list_pods("checkout")
for target in targets:
page = http_get(target.url + "/metrics")
if page is None:
store.write("up", target, 0)
continue
store.write("up", target, 1)
for series, value in parse(page):
store.write(series, target, value)
App.push
each pod sends its own totals when it chooses
def push_loop(target): # the monitor, or a relay
while True:
totals = read_totals()
target.send(self.name, totals)
sleep(push_interval) # 15 s (example)
Relay.serve
keeps what senders sent; answers the monitor's fetch
last = {}
def on_push(sender, totals):
last[sender] = totals
def on_scrape():
return [(sender, totals)
for sender, totals in last.items()]
BatchJob.run
a job that exists only while it works
def run_job(target):
serve("/metrics") # open only while running
result = do_work() # 3 s, 30 s or 5 min
if target is not None:
target.send("batch-job", result)
exit()

Where this sits in Metrics / Monitoring System

Scene 04 of 18, in the Collect act — Running totals, and who starts the conversation.. When the monitor pulls each target on its own schedule, a failed fetch itself says the target is down; when apps push, silence becomes ambiguous — dead, or just quiet?

Up next. Now that the monitor reliably collects every pod's running totals, the next question is how to turn ever-growing totals from many pods into one honest requests-per-second line.

All 18 scenes in Metrics / Monitoring System · Every curriculum

Built with Arqly
Every scene in Metrics / Monitoring System builds on the one before it.All 18 Metrics / Monitoring System scenes