Polling CDC — the lossy fix

Polling the table with WHERE updated_at > :last_seen is technically Change Data Capture, but it cannot see deletes, collapses A→B→A flips between polls, and floors latency at the poll interval.

Previously

Dual-write is broken because there is no commit boundary covering both the DB row and the Kafka event — so the natural reach is for some way to drive the event from the DB itself. Polling the table for changes is the obvious first attempt — does it work?

Scene 02

Polling CDC — the lossy fix

  1. Watch
  2. Try it
  3. Predict
  4. Capture
poll every 5slast_seen = t-0sSELECT WHERE updated_at > :last_seenorders tablehistory per row →r1pendingr2pendingr3pendingr4pendingemitKafka topic(waiting for poll)Polled every 5s · last_seen=0s · 0 events downstream
↓ Change Data Capture — turning DB writes into a stream of change events
What to watch for

We need a way to observe every row change as a stream — a single source feeding many consumers. The first attempt is the path of least resistance: a job that polls the table on an interval and ships whatever has changed.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

PollingJob.pollOnce(last_seen)
the SELECT loop — emits whatever the WHERE clause returns
def pollOnce(last_seen):
rows = db.query(
'SELECT id, status, updated_at FROM orders'
' WHERE updated_at > :last_seen',
last_seen=last_seen,
)
for row in rows:
topic.emit(rowId=row.id, status=row.status)
return max(r.updated_at for r in rows) or last_seen
PollingJob.detectDeletes() # impossible
the function that cannot be written with a SELECT
def detectDeletes():
# A hard-deleted row has no row.
# SELECT ... WHERE updated_at > :last_seen
# cannot return what is not there.
# There is no DELETE-marker in the table.
# There is no updated_at on a non-existent row.
raise NotImplementedError(
'polling CDC cannot observe deletes',
)
Three structural defects
what polling CDC misses, by construction
# 1. missing-deletes:
# no row -> no SELECT match -> silent loss
# 2. collapsed-flips:
# A -> B -> A between polls -> 0 events
# A -> B -> C between polls -> 1 event (C only)
# 3. latency-floor:
# end_to_end_latency >= poll_interval
# shrinking it costs DB load on every poll

Where this sits in Build a CDC pipeline (Debezium + outbox)

Scene 02 of 12. SELECT WHERE updated_at > last_seen is technically Change Data Capture but cannot see deletes, collapses intra-interval flips, and trades latency against DB load.

Up next. Polling cannot be the answer — it loses deletes and collapses fast flips — so we need a source of truth that records every change in order, ideally one the database already keeps.

All 12 scenes in Build a CDC pipeline (Debezium + outbox) · Every curriculum

Built with Arqly
Every scene in Build a CDC pipeline (Debezium + outbox) builds on the one before it.All 12 Build a CDC pipeline (Debezium + outbox) scenes