Submit Order (Prevent Double-Charge)
Worked solution

Submit Order (Prevent Double-Charge) — a worked solution

Idempotency keys, dedup window, retry storms.

Try it yourself first.

You will remember almost none of this if you read it cold. The workspace walks the same 10 stages and runs the architecture you draw through a simulator, so you find out where your version breaks before you see ours.

Open the Submit Order (Prevent Double-Charge) workspace

The problem

Build the submit-order endpoint of a commerce / payments system so that a user is charged exactly once and gets exactly one order — no matter how many times the request is retried. This is a focused problem: we are not designing a whole marketplace (no search, no catalog, no cart, no waiting room). We are designing the one write that touches money, and the three failure realities that make it hard:

  1. The network lies. The response to a successful POST /orders can be dropped before it reaches the client. The client cannot tell "charged" from "never arrived", so it retries — and a naive server charges again.
  2. Two systems can't commit atomically. The charge lives in Stripe; the order lives in your database. There is no distributed transaction across them. A crash between the charge and the order commit leaves the user debited with nothing to show for it.
  3. Failures synchronize. When a downstream (the payment provider) blips, every client retries at once. Without discipline, that 50%-failure moment amplifies traffic ~3.5× and the system cannot recover — a retry storm.

The load-bearing answer is an idempotency key: a client-minted token that names the intent ("this one Pay action"), carried on every retry. The server claims the key in a durable store before doing anything expensive, records its progress as it goes, and on any repeat of the key returns the original outcome instead of redoing the work. Layered on top:

  • a dedup window — how long the key is remembered (24h, aligned with the payment provider's own window),
  • a recovery-point state machine so a half-finished intent can be resumed to completion rather than restarted,
  • and retry-storm hygiene (client backoff + jitter, a gateway retry budget) so retries stay cheap and bounded.

The canonical references are Stripe's idempotency contract and Brandur's Postgres implementation — both are cited throughout.

The reference architecture

Reference architecture for Submit Order (Prevent Double-Charge): 9 components — Client, API Gateway, Service · Order, SQL DB · Orders + Idempotency, External · Payment Provider, Scheduler · Reconciler, Worker · CDC Relay, Stream · Order Events, Worker · Fulfillment — connected by 10 flows.ClientReact (web) + iOS/Andro…API GatewayEnvoy + ext_authzService · OrderGo 1.22SQL DB · Orders + Ide…Postgres 16 (Patroni HA…External · Payment Pr…Stripe + Adyen (dual-pr…Scheduler · ReconcilerKubernetes CronJob + wo…Worker · CDC RelayDebezium (Postgres logi…Stream · Order EventsKafka 3.7 (RF=3, acks=a…Worker · FulfillmentGo + Kafka consumer gro…
9 components, 10 flows. A dashed line is an asynchronous hop. This is the reference design, not the only one that works.

What each component is for

ClientReact (web) + iOS/Android native

Drives the checkout submit. Mints exactly ONE Idempotency-Key — a v4 UUID — the moment the payment sheet OPENS (not at button-press), persists it alongside the in-progress order intent, and re-sends that SAME key on every retry of the same intent. Retries only on network error / timeout / 5xx, using exponential backoff with FULL jitter (sleep = random(0, min(cap, base·2^attempt))), capped at a small attempt budget (~3 attempts in 30s). Renders the server's 201 / 409 / 422 verbatim; never silently re-submits after a 4xx.

Why it exists. The key MUST be minted client-side. The exact failure it defends against — the response to a SUCCESSFUL POST is dropped, so the client cannot tell 'charged' from 'never arrived' — is only survivable if the retry carries the same key. A server-minted key is regenerated on the retry and the dedup fails, producing a second charge. That single fact is why the client is drawn explicitly in this design.

When it fails. User force-quits after the charge but before the 201 renders. On relaunch the client re-sends the same key; order-svc returns the cached 201 (or the reconciler has already completed the order). If the user NEVER returns, the reconciler plus the PSP's 24h cache close the loop — the charge is either completed into an order or refunded. The user is never left charged-with-nothing.

API GatewayEnvoy + ext_authz

Terminates TLS, authenticates the caller, and enforces the idempotency contract at the door: a POST /orders WITHOUT an Idempotency-Key header is rejected 400 before it can reach order-svc (a missing key can never create a duplicate-prone write). Applies per-user rate limits AND a global retry-budget shed — during a downstream brownout it returns 429 Retry-After once aggregate traffic exceeds a budgeted fraction of steady-state, so a retry storm cannot amplify into origin overload. Forwards the key + a correlation id upstream unchanged.

Why it exists. One chokepoint to (a) guarantee every write carries a dedup key and (b) cap retry amplification. Rejected 'each service enforces its own key + limits' — drift lets an un-keyed POST slip through the one service that shipped late, and per-instance rate limits cannot see a global storm. Client-side throttling / retry budgets are the Google SRE 'Handling Overload' prescription; the gateway is where the server half of that lives.

When it fails. A bad rate-limit rule sheds legitimate checkout traffic (directly revenue-affecting). Mitigation: progressive config rollout with auto-rollback on 5xx > 0.5% / 60s; the shed is a 429 with Retry-After (client backs off, key preserved) — NOT a hard drop — so a mis-tuned budget delays orders rather than double-charging or losing them. Alarm on gw.429_rate decoupled from downstream health, so 'we are shedding good traffic' is distinguishable from 'downstream is actually down'.

Service · OrderGo 1.22

The idempotent receiver and orchestrator for POST /orders. Runs Brandur's recovery-point state machine over the Orders DB: (1) in its OWN transaction, UPSERT the idempotency_keys row with ON CONFLICT — a fresh row is created at recovery_point='started' with request_hash + locked_at stamped; a losing INSERT loads the existing row. (2) If the row is 'finished', return the cached response_code / response_body immediately — no charge, no order. (3) If locked_at is still fresh (a concurrent sibling holds it), return 409 in-progress. (4) Otherwise RESUME from the row's recovery_point. The phases alternate LOCAL atomic phases with the SINGLE foreign state mutation (the charge): started → [charge the Payment Provider, forwarding the key] → charge_created → [SERIALIZABLE txn: INSERT order + INSERT outbox + UPDATE the idempotency row to 'finished' with the cached response] → finished. Each phase advances recovery_point in the SAME txn, so a crash anywhere is restart-safe.

Why it exists. There must be exactly ONE component that owns the atomic boundary between 'money moved' and 'order recorded', with dedup on top. Rejected 'client calls Stripe then calls our DB' — the user could be charged with no order and no compensating undo. Rejected 'dedup in Redis, then charge, then write DB' — a Redis-only dedup record is lost on eviction / OOM and the retry double-charges. The recovery-point pattern makes every partial failure drivable to completion by the next retry OR the reconciler — that is the whole game.

When it fails. Crash between charge_created and finished — provider charged, order not yet written. Recovery: the client's next retry loads recovery_point='charge_created', re-calls the provider (cached charge, same key → no second charge), completes the commit, returns 201. If the client never retries, the reconciler sweeps the stuck row. Blast radius of losing one order-svc replica is ZERO orders — it is stateless; all durable state lives in Postgres.

SQL DB · Orders + IdempotencyPostgres 16 (Patroni HA, sync replica)

System of record AND the dedup authority. Three tables in one shard so they commit together: orders (durable order + charge_id, with a UNIQUE dup guard), idempotency_keys (client key + request_hash + recovery_point + locked_at + cached response_code / response_body; unique on (user_id, key)), and outbox (side-effect events published post-commit by CDC). Sharded by hash(user_id) so one user's retries AND their order always land on the SAME shard — the dedup claim and the order commit are same-shard, no cross-shard transaction. idempotency_keys is declaratively partitioned by key_date (derived from the key) with a daily DROP PARTITION reaper.

Why it exists. The dedup record must be as durable as the order itself and committed in the SAME transaction as order + outbox, or the guarantee has a gap. Postgres gives ACID across all three tables in one txn, a UNIQUE constraint as the last-line dup defense, and SERIALIZABLE isolation. Rejected DynamoDB conditional-writes (workable, but the order+outbox+idempotency atomicity wants a real multi-row txn — TransactWriteItems caps at 100 items / 4MB and complicates the outbox); rejected Redis-as-dedup (not durable — the precise anti-pattern this problem exists to teach).

When it fails. Primary loss = ~30s Patroni failover write outage; idempotency makes in-flight retries safe post-promotion (client re-sends the same key, hits the promoted primary, resumes). RPO=0 via the sync replica so no committed order / charge is lost. The dangerous slow-burn is the partition-drop reaper failing — idempotency_keys bloats and the unique-index probe on the claim slows every confirm; page on idempotency_keys.size_gb per shard. A too-SHORT dedup window is a config error (deep-dive #3), not a runtime failure.

External · Payment ProviderStripe + Adyen (dual-provider abstraction)

Moves the money. order-svc calls POST /charges (PaymentIntents) with Idempotency-Key: <the same client key, forwarded>. Stripe persists the first response (status + body) under that key and REPLAYS it for 24h — so order-svc's crash-recovery resume re-issues the IDENTICAL call and gets the original charge back instead of a second charge. Returns charge_id on success; order-svc attaches it to the order row. Concurrent requests with the same key at Stripe get a 409 (not cached, safe to retry); the same key with different params gets an error (422).

Why it exists. We do not move money ourselves — staying PCI SAQ-A means never touching raw card data, only PSP tokens. Stripe's own idempotency layer is the SECOND line of double-charge defense beneath ours: even if our recovery point re-drives the charge, the forwarded key makes Stripe dedup it. Dual-provider (warm Adyen) is the mitigation for a single-provider outage (the Visa Europe 2018-class event) so a Stripe brownout degrades to Adyen rather than stranding charges.

When it fails. Stripe hard outage — detected by provider error_rate > 50% / 60s → route to Adyen; the idempotency-key semantics are normalized across both by the abstraction layer. If BOTH providers are down, order-svc fails the confirm with a 503 (key preserved, NO charge attempted) and the client backs off — so no order is ever created without a successful charge, and no charge is ever stranded without an order.

Scheduler · ReconcilerKubernetes CronJob + worker pool

The backstop for intents abandoned mid-flight (charged, then the client never retried). Every 1 minute, per shard: SELECT idempotency_keys rows WHERE recovery_point <> 'finished' AND locked_at < now() - lock_timeout ... FOR UPDATE SKIP LOCKED. For each: if recovery_point='charge_created' (money moved, order not written) → drive the commit to completion (write order + outbox, mark finished); if the order can no longer be honored (item gone, price changed) → issue a refund via the Payment Provider (refund-scoped key) and mark the row terminal; if 'started' past a grace window with no charge → expire it.

Why it exists. The recovery-point resume only fires if SOMEONE retries. A user who charged and closed the tab forever needs an automated closer, or you get charged-but-no-order sitting silently — the worst possible outcome for trust. Rejected 'rely on the PSP webhook alone' — webhooks can be delayed / lost and do not cover the 'started but never charged' rows; the reconciler is the authoritative sweep over OUR own durable state. Invariant it guarantees: no idempotency row stays non-'finished' past a bounded window without either a completed order or a refund.

When it fails. Reconciler down → charged-but-no-order rows linger until it resumes (bounded, not lost — the durable rows ARE the work queue). The dangerous case is a reconciler BUG that refunds an order that was legitimately completing: guarded by lock_timeout > confirm-p99 (never touch in-flight), FOR UPDATE row locks, and an alarm on refund_rate > baseline. Its worst failure is latency (delayed closure), never a lost or duplicated charge.

Worker · CDC RelayDebezium (Postgres logical replication)

Active-passive Debezium sidecar tailing each shard's logical replication slot, filtered to the outbox table (outbox-event-router SMT). For each COMMITTED outbox row it publishes to Kafka (order.confirmed / order.canceled / order.refunded) keyed by order_id, then advances the slot offset. The order commit and the outbox row are ONE Postgres transaction, so the event is published exactly-once-at-the-source; the relay is at-least-once INTO Kafka and downstream consumers are idempotent.

Why it exists. To tie the order commit to the downstream publish WITHOUT a dual-write. If order-svc wrote Postgres AND published to Kafka directly, a crash between the two diverges permanently — an order with no receipt, or a receipt for an order that rolled back. The outbox + CDC pattern is the only known-correct way (deep-dive #4). Without a relay node the orders-db → bus edge would be a fiction — Postgres does not call Kafka on its own.

When it fails. Relay stalls → the replication slot grows → primary pg_wal disk pressure → writes threatened. Page on replication_slot_lag_bytes well before disk fills; auto-restart with offset recovery from the last published row. This does NOT affect the double-charge guarantee (that is all committed in Postgres already) — it only delays side-effects.

Stream · Order EventsKafka 3.7 (RF=3, acks=all, ISR=2)

Durable event bus carrying order.confirmed / order.canceled / order.refunded, published by the CDC relay and consumed by the fulfillment worker. Partitioned by order_id (per-order ordering), RF=3 with acks=all and min.insync.replicas=2, 7-day retention for replay after a consumer bug.

Why it exists. Decouples the user-blocking confirm from the side-effects (receipt, fulfillment) so a slow email provider cannot inflate confirm p99, and gives each consumer its own retry budget + replay. Rejected fire-and-forget straight from order-svc — that loses the durability and replay the outbox buys us, and re-introduces the dual-write divergence.

When it fails. Broker loss → leader re-election ~5s; alarms on under_replicated_partitions and order.confirmed publish_age_p99 (the latter is the business pager — 'user paid but no receipt for N minutes'). A consumer lag spike delays receipts / fulfillment but never affects charge / order correctness — those are already committed.

Worker · FulfillmentGo + Kafka consumer group

Consumes order.confirmed and performs the order's side-effects exactly-once-effectively: send the receipt / confirmation email + push, trigger fulfillment (warehouse pick or digital entitlement), decrement inventory, then UPDATE the order row's fulfilled_at. Every downstream call carries order_id as its OWN dedup key, so a redelivered Kafka message (at-least-once) does not send two receipts or ship twice.

Why it exists. The submit-order guarantee is not only 'one charge' — it is 'one receipt, one fulfillment'. At-least-once delivery from Kafka means the consumer MUST be idempotent or a redelivery double-emails / double-ships. This is the SAME dedup discipline as the charge, applied to side-effects — which is why it is in scope here and not a bolt-on feature.

When it fails. A poison message stalls a partition → bounded retries then DLQ + page; a non-idempotent handler bug would double-send receipts — guarded by the order_id dedup and the conditional fulfilled_at write. Never touches the charge / order rows, so it can never cause a double-charge.

Stage by stage

The same 10 stages the workspace walks, answered.

01Clarifications

What would you ask before drawing a single box?

Typical clarifications to surface:

  • Is the charge synchronous or does it settle async? Assume a synchronous authorize+capture against a PSP (Stripe/Adyen) on the request path — that is the hard case (a foreign state mutation inside the request). Async settlement only moves the same problem to a webhook.
  • Who mints the idempotency key — client or server? Client. This is not a style choice: the retry we must survive is the one where the client never saw the first response, so only the client can re-supply the same key. A server-minted key regenerates and the dedup fails.
  • How long must a key be honored (the dedup window)? Long enough to cover the client's entire retry horizon (including an app relaunch hours later) and aligned with the PSP's window. Default 24h (Stripe's).
  • Exactly-once, or effectively-once? There is no true exactly-once across two systems. We guarantee at-most-one charge + at-most-one order, and at-least-once side-effects with idempotent consumers (effectively-once).
  • What is the natural key of an order? Is a duplicate defined by the idempotency key alone, or also by (user, cart-hash)? We use the idempotency key as the dedup identity and keep a request fingerprint to reject key-reuse-with-different-body.
  • Single region? Yes — the money tier is CP and single-region here. Multi-region residency is out of scope for this focused problem.

Assumptions to state:

  • 5M orders/day, ~58/s average, 20× flash-sale peak (~1,160 orders/s).
  • Payment provider p99 ~2s; dual-provider (Stripe primary, Adyen fallback).
  • Dedup window 24h; keys are v4 UUIDs minted at payment-sheet open.
  • Client retries: full-jitter exponential backoff, ≤3 attempts in 30s.

02Functional reqs

What must this system actually do?

  • Submit an order (POST /orders) with a client-minted Idempotency-Key. The server authorizes+captures payment and records the order. Must be idempotent: any number of retries of the same key ⇒ one charge, one order, one response.
  • Replay safety: a repeated key returns the original response (status + body), including a repeated error, without re-executing.
  • Concurrent-duplicate safety: two in-flight requests with the same key ⇒ one proceeds, the other gets 409 in-progress (retriable), never a second charge.
  • Resume a half-finished intent: if the process crashed after charging but before recording the order, the next retry (or the reconciler) drives it to completion — no second charge.
  • Reject key misuse: the same key with a different request body ⇒ 422, never the wrong cached order.
  • Cancel / refund an order idempotently (same key discipline, refund-scoped provider key).
  • Deliver side-effects exactly-once-effectively: receipt email, fulfillment, inventory decrement — driven from the committed order via an outbox, with idempotent consumers.

Out of scope (deliberately): catalog, search, cart, recommendations, fraud scoring, tax/shipping calculation, a waiting room. This problem is only the double-charge-safe submit.

03Non-functional

What must it promise about speed, uptime and correctness?

  • Correctness (the top invariant): at-most-one charge and at-most-one order per idempotency key, under arbitrary retries, crashes, and a single-node failure at any step. This dominates every other property.
  • Availability: 99.9% on the confirm path. A brief write outage (failover) is acceptable because idempotency makes the client's retry safe.
  • Latency: confirm p99 5–10s (payment-provider-dominated); an idempotent replay p99 < 50ms (one indexed row read); concurrent-duplicate rejection < 30ms.
  • Durability: RPO=0 on orders and charges (sync replication) — a lost committed order is a real double-charge exposure on the retry. Orders retained 7+ years (financial). Idempotency keys retained for the 24h dedup window.
  • Consistency: the order + idempotency-response + outbox commit is ONE SERIALIZABLE transaction. The dedup authority is the SQL DB, never a cache.
  • Overload behavior: a retry storm must be shed, not passed through — accepted traffic bounded to ~1.2× steady-state; excess returns 429 Retry-After (key preserved).

04Capacity estimation

How much load and data does this have to hold?

Anchor: a large merchant on a flash-sale / Black-Friday peak, where checkout is bursty and retries spike exactly when the payment provider is stressed.

Per-tier QPS (formulas, then numbers — 5M orders/day, 20× peak, 1.3 attempts/order):

  • Order rate: ordersPerDay / 86400 ≈ 58/s avg; at 20× peak ≈ 1,160 orders/s.
  • Submit attempts (API-facing, incl. retries): ordersQpsPeak × retriesPerOrder ≈ ~1,500/s at peak. This is what the gateway + order-svc actually receive.
  • Retry-storm ceiling (the shed target): ordersQpsPeak × retryStormMultiplier ≈ ~5,800/s worst case during a downstream blip. The gateway's retry budget caps accepted traffic near 1.2× steady (~1,800/s) and sheds the rest — this number is the thing we are defending against, not provisioning for.
  • Payment charges (deduped): equals the order rate — ~1,160/s at peak. Dedup means retries do NOT multiply provider calls.
  • In-flight charges (Little's law): chargesQpsPeak × p99_seconds ≈ 1,160 × 2 ≈ ~2,320 concurrent provider calls — the connection-pool + timeout budget must cover this.
  • Async side-effects: ordersQpsPeak × fanoutPerOrder ≈ ~4,640 events/s through Kafka. Trivial at RF=3.

Bytes economics — explicitly checked: every payload here is JSON < 10KB (an order, a charge, an event). There is no large-body / direct-to-storage concern — do NOT import upload/CDN patterns from image or video canonicals. Service tiers correctly sit in the data path.

Storage:

  • Idempotency keys (24h window): ordersPerDay × 1KB ≈ ~5 GB for the window (key + request_hash + cached response). Across 8 shards, sub-GB each. With two live daily partitions, ~10 GB/shard ceiling.
  • Orders (7yr): ordersPerDay × 365 × 7 × 2KB ≈ ~26 TB raw across 8 sharded primaries. Fine.
  • Order events (Kafka, 7-day retention): ~4,640/s × 1KB × 7d ≈ ~2.8 TB — trivial for a 3-broker cluster.

Architectural flip points (where the shape changes, not just the numbers):

  • Redis-only dedup breaks the instant it evicts. A cache with a TTL and an eviction policy will, under memory pressure, silently drop a key that a delayed retry still references — and double-charge. The dedup authority MUST be durable (SQL). Redis is legitimate only as an optional fast in-flight lock, never as the record of truth.
  • The dedup window must be ≥ the client's max retry horizon. If a key can be reaped before the last possible retry arrives (e.g. an app relaunched next morning), you get a second charge. Sizing the window is a correctness knob, not a cost knob — align it with the PSP (24h).
  • Single-shard commit is a hard requirement. Sharding by hash(user_id) keeps a user's retries and their order on one shard so the claim and the order commit are one local transaction. If order and idempotency ever land on different shards, you have re-created the cross-store dual-write you were trying to avoid.

05API design

What does the outside world call, and what comes back?

POST /orders
Content-Type: application/json
Authorization: Bearer <session>
Idempotency-Key: 73c8e8e0-3a2d-4f5b-8d2e-1c2b3a4d5e6f   # client-minted v4 UUID, one per Pay intent

{
  "cart_id": "c_9931",
  "amount_cents": 48200,
  "currency": "usd",
  "payment_token": "pm_card_visa_xxxx"
}

201 Created
{ "order_id": "o_ab12cd", "charge_id": "ch_3Pxyz...", "status": "confirmed" }

# --- the same key, replayed after a dropped response: returns the ORIGINAL 201, no re-charge ---

409 Conflict
{ "error": "request_in_progress" }   # a sibling with the same key is mid-flight; retry after a short delay

422 Unprocessable Entity
{ "error": "idempotency_key_reused_with_different_body" }   # request fingerprint mismatch

504 Gateway Timeout + Retry-After
# order-svc contract: client retries on 504 with the SAME key → resumes / returns cached 201 within 24h

429 Too Many Requests + Retry-After
# retry-budget shed at the gateway during a storm — key preserved, back off and retry
POST /orders/:order_id/cancel
Idempotency-Key: 8a2f4c6e-1b3d-4e5f-9a7b-2c4d6e8f0a1b

200 { "order_id": "o_ab12cd", "status": "canceled", "refund_status": "pending" }

Response-caching rule (matches Stripe): the first request's status + body is cached for the key and replayed for both success and error (including a 500). Not cached: validation failures (400/422) and 409 concurrent-conflicts — those never began execution, so they are safely retriable rather than sticky.

06Data model

What gets stored, and what is it looked up by?

idempotency_keys (Postgres, sharded by hash(user_id), partitioned by key_date) — Brandur's shape:

fieldtypenotes
keyuuidclient-minted v4 UUID (≤ 255 chars if opaque string)
key_datedatepartition column; derived from the key's day of first use
user_idbigintkey is scoped per-user — a stolen key can't cross accounts
request_hashbyteaSHA-256 of the canonicalized body; mismatch on reuse ⇒ 422
recovery_pointtextstarted → charge_created → finished
locked_attimestampset while a request holds it; a fresh lock ⇒ 409 in-progress
response_codeintcached status of the first request
response_bodyjsonbcached body of the first request
created_attimestamp

PK (user_id, key, key_date). key_date is a pure function of the key, so a retry always targets the same live partition and still conflicts on insert. A daily DROP PARTITION reaper bounds the table to the dedup window (two live partitions cover 24h across midnight).

orders (Postgres, same shard, same txn):

fieldtypenotes
order_iduuidPK
user_idbigintshard key
idem_keyuuidthe key that created it (1:1)
charge_idtextfrom the Payment Provider
statusenumconfirmed / canceled / refunded
amount_centsbigint
fulfilled_attimestampset by the idempotent fulfillment consumer (WHERE NULL)
created_attimestamp

UNIQUE (idem_key) — the belt-and-braces guarantee that one key ⇒ at most one order row, even if the application logic is bypassed.

outbox (Postgres, same shard, same txn): (outbox_id, topic, payload jsonb, created_at) — the CDC relay reads here and publishes to Kafka, then advances the WAL slot.

Database choice — recommended: sharded Postgres. ACID across orders + idempotency_keys + outbox in ONE transaction; a UNIQUE constraint as the last-line dup defense; SERIALIZABLE isolation for the commit phase; mature ops tooling. Why not Redis for dedup? It is not durable — eviction under memory pressure drops a key a delayed retry still needs, and the retry double-charges. Why not DynamoDB? Conditional writes work for the claim, but the order+outbox+idempotency atomicity wants a real multi-row transaction; TransactWriteItems caps at 100 items / 4MB and complicates the outbox. Why not Spanner/CockroachDB? Fine if you need geo-distributed serializable writes, but +50% write latency and 3× cost — the commit is single-shard by construction, so the cost isn't warranted.

07High-level design

Which components handle a request, and in what order?

The load-bearing flow (submit):

  1. Client mints one Idempotency-Key at payment-sheet open and re-sends it on every retry (full-jitter backoff, capped attempts).
  2. API Gateway rejects a keyless POST /orders (400), rate-limits per user, and applies a retry budget — during a storm it sheds excess as 429 Retry-After so amplification can't reach origin. It forwards the key unchanged.
  3. Order Service runs the recovery-point state machine over Postgres:
  • Claim the key in its own txn (ON CONFLICT DO NOTHING RETURNING). If the row is finished, return the cached response — done, no charge. If it's freshly locked_at, return 409. Else resume from its recovery_point.
  • started → charge the Payment Provider, forwarding the key (the one foreign state mutation) → advance to charge_created.
  • charge_created → one SERIALIZABLE txn: INSERT order + INSERT outbox + UPDATE the idempotency row to finished with the cached response → return 201.
  1. CDC Relay (Debezium) tails the WAL, publishes order.confirmed from the outbox to Kafka; the Fulfillment Worker consumes it and does receipt + fulfillment idempotently (keyed on order_id).
  2. Reconciler sweeps any idempotency row stuck below finished past its lock timeout and drives it to completion or refunds it — the backstop for a client that charged and never came back.

Data flow on the first successful submit: Client → Gateway → Order Service → (claim started) → Payment Provider (charge) → (SERIALIZABLE commit: order + outbox + finished) → 201. Async tail: WAL → CDC → Kafka → Fulfillment.

Data flow on a duplicate retry: Client → Gateway → Order Service → orders-db (row is finished) → return the cached 201. No charge, no order, ~50ms.

Data flow on a concurrent duplicate: two requests race the claim; the ON CONFLICT loser (or a fresh locked_at) sees the row locked → 409 in-progress; the client retries after a short delay and then hits the finished path.

Data flow on crash recovery: Order Service died at charge_created. Next retry loads the row, re-calls the provider with the same key (provider replays the original charge — no second charge), commits the order, marks finished, returns 201.

Multi-region posture (scoped out but stated): single-region CP money tier. If residency forces a second region, partition orders by region-of-origin with no cross-region locks; the idempotency claim stays local to the order's home region.

08Deep dives

Which part breaks first, and what do you do about it?

1. Double-charge under a retry storm (the headline). User mashes "Pay" 3× in 200ms. All three carry the same client key (minted at sheet-open). The Postgres ON CONFLICT claim serializes them: the first creates the started row and proceeds; the second and third either find a fresh locked_at (→ 409, retriable) or, once the first finishes, find recovery_point='finished' and return the cached 201. Exactly one charge, one order. The naive path that double-charges: charging the provider before claiming the key (all three charge), or deduping in Redis (an eviction between attempts drops the record and the retry charges again). The rule: claim in a durable store before the foreign mutation; the foreign mutation forwards the same key so the provider dedups too.

2. Charged-but-no-order (the atomicity gap). There is no transaction across Stripe and Postgres. If the provider charges and then the commit fails (deadlock, primary failover), the row is stuck at charge_created. Recovery is built into the state machine: the next retry resumes from charge_created, re-calls the provider with the same key (which replays the original charge for 24h — no second charge), commits the order, and marks finished. If the client never retries, the reconciler sweeps the stuck row (older than the lock timeout) and completes it — or refunds via a refund-scoped key if the order can no longer be honored. The invariant: no key stays below finished past a bounded window without either an order or a refund.

3. Dedup window sizing (the subtle correctness knob). The key is remembered for a window (24h). If a retry arrives after the key is reaped — an app relaunched the next morning, a delayed mobile retry — the server treats it as brand-new and charges again. So the window must be ≥ the client's maximum retry horizon and ≥ the payment provider's own window (Stripe = 24h), and the two should be aligned: if Stripe forgets the charge key before you forget yours (or vice-versa), the "provider dedups my resume" guarantee breaks. Bound growth with a daily partition drop; because key_date is derived from the key, a retry inside the window always lands in a live partition. Too short ⇒ double-charge; too long ⇒ table bloat and a slow claim (see failure scenario idempotency-table-bloat).

4. Same key, different body (misuse + attack). A client bug (or an attacker replaying a captured key) sends the same key with a different cart. If the server blindly returned the cached response, the second user would get the first user's order — a correctness and security hole. Defense: store a request fingerprint (SHA-256 of the canonicalized body) on claim; on reuse, compare. Match ⇒ serve the cached response. Mismatch ⇒ 422, and never execute or serve the wrong order. The key is scoped per-user (unique (user_id, key)), so a stolen key cannot cross accounts. This mirrors Stripe returning an error when parameters don't match the original request.

5. Retry-storm collapse (keeping retries cheap and bounded). The provider 503s for 60s; every client retries; the API sees ~5× amplification. Three layers keep it from melting: (a) client full-jitter backoff (sleep = random(0, min(cap, base·2^n))) with a hard attempt cap — full jitter flattens the synchronized spike that plain backoff leaves; (b) gateway retry budget — accept ~1.2× steady-state and shed the rest as 429 Retry-After, so a 50%-failure self-amplification (~3.5×) can't pass through and prevent recovery (Google SRE "Handling Overload"); (c) the idempotent replay is cheap — a repeated key is one indexed row read, not a re-charge, so even the retries that do get through cost almost nothing. Add a circuit breaker + dual provider so recovery is graceful. The elegance: idempotency and retry-budgeting reinforce each other — because replays are safe and cheap, aggressive client retries are tolerable; because retries are budgeted, the provider gets room to recover.

09Trade-offs

What did this design cost, and what breaks at 10×?

What we accepted (vs. the stronger alternative):

  • Effectively-once, not exactly-once. True exactly-once across Stripe + Postgres is impossible (no distributed txn). We guarantee at-most-one-charge + at-most-one-order + at-least-once idempotent side-effects. The residual "charged-but-no-order" window is closed by recovery-point resume + reconciler, not by a magic 2PC.
  • A synchronous charge on the request path (confirm p99 5–10s). The alternative — accept the order, charge async — trades a slow confirm for a "your order may fail after the fact" UX and a webhook that re-introduces the same idempotency problem. For this focused problem the synchronous charge is the honest hard case.
  • Single-region money tier. Multi-region active-active writes need cross-region locks or CRDTs on the order/charge — prohibitive. A region failover is a ~30s write outage; idempotency makes the client's retry safe across it.
  • Dedup authority in SQL, not Redis. We pay a Postgres round-trip on the claim (vs. a sub-ms Redis SET) to get durability. The whole problem exists because the fast-but-lossy option double-charges; the latency is worth it.
  • Async side-effects (receipt/fulfillment) via the outbox. There's a ~5s gap between 201 and the receipt email. Acceptable: the user has the order_id in the response; the email is a side-effect delivered at-least-once by an idempotent consumer.

What breaks at 10× scale (hypothetical ~12K orders/s sustained):

  • Single-shard claim contention — hot users or a mis-chosen shard key concentrate claims; move to more shards, keep hash(user_id) so one intent stays local.
  • SERIALIZABLE serialization failures (40001) rise under contention — bounded application retries with backoff; keep the commit txn short (three inserts, no external calls inside it — the charge is outside the txn by construction).
  • Idempotency partition churn — daily drops become large; move to hourly partitions or a shorter window if (and only if) it still covers the client + PSP retry horizons.
  • CDC relay throughput — Debezium ~10K events/s/replica; at 12K orders × 4 fan-out = 48K events/s you need ~5 relays, or a Kafka-native outbox log.

Trace catalogue

The simulator authors 6 problem-specific traces — the journeys a senior engineer expects to walk when reviewing this design. The load-bearing insight across the first four: the path is nearly identical (Client → Gateway → Order Service → Orders DB / Payment Provider); what differs is the durable state the idempotency claim finds — that is idempotency in one sentence:

  • submit-order:happy-path-first-attempt — the baseline write. Client → Gateway → Order Service → (claim started) → Payment Provider (charge) → (SERIALIZABLE order + outbox + finished) → 201. Budget ~2.5s (provider-dominated).
  • submit-order:duplicate-retry-replay — the load-bearing dedup. A repeated key: Client → Gateway → Order Service → orders-db (row finished) → return the cached 201, no re-charge. Budget ~50ms.
  • submit-order:concurrent-duplicate-inflight — two identical requests race the claim; one proceeds, the other sees a fresh locked_at → 409 in-progress; client backs off. Budget ~30ms.
  • submit-order:crash-recovery-resume — Order Service died at charge_created. The retry resumes: load row → re-call provider (cached charge, same key, no second charge) → commit order → finished. Budget ~1.5s. (The same-key-different-body → 422 fingerprint case ships as the key-reuse-wrong-body failure scenario below.)
  • submit-order:async-fulfillment-outbox — the exactly-once-effectively side-effect. orders-db WAL → CDC → Kafka → Fulfillment Worker (idempotent receipt + fulfilled_at WHERE NULL). Budget ~5s.
  • submit-order:reconcile-charged-no-order — the abandoned-intent backstop. Reconciler sweeps a stuck charge_created row (client never retried) → completes the order, or refunds via a refund-scoped key. Budget ~60s.

Failure scenarios we model

6 chaos scenarios, each grounded in a cited real-world precedent:

  • double-submit-network-blip (process) — the response to a successful POST is dropped; the client retries; the design must produce one charge and one order. Stripe's idempotency rationale is the canonical precedent.
  • payment-provider-blip-retry-storm (deps) — provider 503 for 60s → client retry amplification ~5× → gateway retry-budget shed + client full-jitter backoff hold the line (AWS "Timeouts, retries, and backoff with jitter"; Google SRE "Handling Overload").
  • charged-but-no-order-commit (data) — provider charged, then the Postgres commit failed (deadlock / primary failover). Recovery-point resume + reconciler close the gap with no second charge (Brandur recovery points).
  • idempotency-window-expiry-double-charge (process) — a delayed retry arrives after the 24h dedup window reaped the key → a second charge. Mitigation: window ≥ max client retry horizon AND aligned with the PSP's 24h window.
  • key-reuse-wrong-body (correctness) — a key is reused for a different cart (bug or attacker replay); the request-fingerprint mismatch must return 422 and never serve the first order — mirrors Stripe's parameter-mismatch error.
  • idempotency-table-bloat (process) — the daily partition-drop reaper fails for days; idempotency_keys grows unbounded; the unique-index probe on the claim slows every confirm. Page on idempotency_keys.size_gb per shard.

Primary sources

  • Stripe — Designing robust APIs with idempotency
  • Stripe API Reference — Idempotent Requests (24h window, 409 concurrent, param-mismatch error)
  • Brandur — Implementing Stripe-like Idempotency Keys in Postgres (recovery points + atomic phases)
  • AWS Builders' Library — Timeouts, retries, and backoff with jitter (Marc Brooker)
  • Marc Brooker — Exponential Backoff And Jitter
  • Google SRE Book — Handling Overload + Addressing Cascading Failures (client-side throttling, retry budgets)

Now defend it

Reading a design is not the same as being able to hold one under questioning. The workspace asks the same questions an interviewer would, and the simulator disagrees with you when the diagram does not support the claim.

Work Submit Order (Prevent Double-Charge) yourself