The dual-write trap — no atomicity boundary across DB and Kafka

A service that commits to its DB and then publishes to Kafka has no atomicity boundary spanning both, so every crash window between the two is a permanent inconsistency the system cannot heal on its own.

Scene 01

The dual-write trap

  1. Watch
  2. Try it
  3. Predict
  4. Capture
Crash windowT1T2T3T4atomicity boundary (DB only)Serviceuser.balance = 100two writes, no shared txnCOMMITDB✓publishKafka topic — events✓Healthy baselineDB says:balance = 100Kafka says:BalanceUpdated(100)Both legs agree — for now.Pick a crash marker to inject a failure.Healthy baseline — both legs succeed. Watch what happens when the timeline catches a crash.
What to watch for

Watch the healthy baseline. The service writes user.balance=100 to its DB, then publishes BalanceUpdated to Kafka — two separate writes to two separate systems. That pair-without-a-shared-transaction is the dual-write problem. As long as nothing crashes, both downstream views agree.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Service.handleSignup(event)
the dual-write — two writes, no shared transaction
def handleSignup(event):
tx = db.begin()
tx.execute('UPDATE users SET balance=100 ...')
tx.commit() # leg 1: DB
kafka.publish('BalanceUpdated', # leg 2: Kafka
{balance: 100})
return ok
Service.handleSignup_with_retry(event)
the 'obvious' fix — wrap publish in retry, walk into T3
def handleSignup(event):
tx = db.begin()
tx.execute('UPDATE users SET balance=100 ...')
tx.commit()
for attempt in range(MAX_RETRIES):
try:
kafka.publish('BalanceUpdated',
{balance: 100})
return ok
except AckTimeout:
continue # broker may have it already
raise PublishFailed
// Four crash windows, one root cause
no atomicity boundary spans both legs
# T1: commit OK, publish dies -> event lost
# T2: publish OK, commit fails -> phantom event
# T3: both OK, ack lost, retry -> duplicate event
# T4: two writers race -> DB and Kafka
# disagree on order
#
# All four are the same bug: the service has one
# transaction (the DB tx); the publish escapes it.

Where this sits in Build a CDC pipeline (Debezium + outbox)

Scene 01 of 12. Service writes to its DB and publishes to Kafka — and any crash between those two writes is permanent inconsistency. Four scenarios, four divergences, one structural fix.

Up next. Dual-write is broken because there is no commit boundary covering both the DB row and the Kafka event — so the natural reach is for some way to drive the event from the DB itself.

All 12 scenes in Build a CDC pipeline (Debezium + outbox) · Every curriculum

Built with Arqly
Every scene in Build a CDC pipeline (Debezium + outbox) builds on the one before it.All 12 Build a CDC pipeline (Debezium + outbox) scenes