Ticketmaster / Hotel Booking
Worked solution

Ticketmaster / Hotel Booking — a worked solution

Don't oversell. Hold-then-confirm.

Try it yourself first.

You will remember almost none of this if you read it cold. The workspace walks the same 10 stages and runs the architecture you draw through a simulator, so you find out where your version breaks before you see ours.

Open the Ticketmaster / Hotel Booking workspace

The problem

Build a system that sells finite inventory under high contention without overselling. The canonical examples are Ticketmaster (concert tickets) and Booking.com / Airbnb / Expedia (hotel rooms) — but the pattern generalizes to airline seats, restaurant tables, exam slots, ride dispatching, any limited-quantity-for-sale resource. The load-bearing technical reality is the same in all of them:

  1. Inventory is finite — there are N tickets to this concert, M rooms at this hotel on this date.
  2. Demand can dwarf supply — Taylor Swift Eras Tour onsale: 14M users showed up for a system sized for 1.5M; 3.5M Verified-Fan pre-registrations alone exceeded the previous all-time peak by 4×.
  3. You can't oversell — selling the same seat twice destroys customer trust and triggers chargebacks / regulatory complaints / class actions.
  4. You can't lose money on retries — when the user mashes "Pay" three times in 200ms because the network blipped, they must end up with one charge and one booking — not three of each, not zero of either.

The pattern that emerged across Ticketmaster, SeatGeek, StubHub, Booking.com, Airbnb, Expedia, and the broader ticketing/lodging industry is hold-then-confirm:

  • User picks a seat / room → system places a soft hold with a TTL (5–15 min) — Redis SET NX PX is the dominant primitive.
  • User completes payment within the hold window → system confirms via a strongly-consistent durable write that ties the booking commit, the idempotency claim, and the outbox publish into one transaction.
  • Hold expires (TTL or explicit cancel) → seat returns to the pool automatically — no janitor job; Redis PX expiry does the work.

The reference architecture

Reference architecture for Ticketmaster / Hotel Booking: 19 components — Client, CDN · Bot Edge, Service · Waiting Room, KV Store · Waiting Room Tokens, API Gateway, Service · Search, Search, Service · Hold, Cache · Holds, Service · Confirm, SQL DB · Inventory + Idempotency, External · Payment Provider, Worker · CDC Relay, Stream · Event Bus, Worker · Fan-out, Service · Fraud Scoring, Worker · Compensation (Saga), Service · PSP Webhook Receiver, Object Store · Audit Sink — connected by 26 flows.ClientReact (web) + iOS/Andro…CDN · Bot EdgeFastly + Cloudflare Bot…Service · Waiting RoomAWS Lambda + API Gatewa…KV Store · Waiting Ro…DynamoDB (on-demand) wi…API GatewayEnvoy + custom auth fil…Service · SearchJava + Spring + OpenSea…SearchElasticsearch 8 (manage…Service · HoldGo + Redis client (redi…Cache · HoldsRedis 7 (cluster mode)Service · ConfirmJava + Spring + Stripe …SQL DB · Inventory + …Postgres 16 (Patroni + …External · Payment Pr…Stripe + Adyen (dual-pr…Worker · CDC RelayDebezium + Kafka ConnectStream · Event BusKafka 3.7 (Confluent / …Worker · Fan-outGo + Kafka consumer gro…Service · Fraud Scori…Java + Sift SDK + Strip…Worker · Compensation…Go + Kafka consumer + T…Service · PSP Webhook…Go + HMAC verifier + Po…Object Store · Audit …S3 + Glue catalog + Par…
19 components, 26 flows. A dashed line is an asynchronous hop. This is the reference design, not the only one that works.

What each component is for

ClientReact (web) + iOS/Android native

Serves three flows: (a) browse / search results, (b) hold a seat / room for N minutes, (c) confirm + pay. Mints the client-side idempotency key (a UUIDv7 per Pay intent — its embedded timestamp derives the key's daily partition in idempotency_keys) and re-sends it on every retry of the same intent. Holds the waiting-room token in localStorage / Keychain and presents it on every subsequent call.

Why it exists. It's the user. Listed explicitly because the client owns the idempotency key — the alternative (server-minted key) breaks on the exact retry it's meant to defend against: when the user mashes 'Pay' and the response packet is dropped, only the client can supply the same key on retry; a server-minted key would be regenerated and the dedup fails.

When it fails. User loses connectivity mid-confirm. Client retries with the same Idempotency-Key when the network returns; payment-svc serves the cached response. If the user kills the app, the hold expires server-side; on next launch the client sees 409 hold_expired and routes back to the seat picker.

CDN · Bot EdgeFastly + Cloudflare Bot Management

Public IP for every request. Caches search/static (TTL 60s with stale-while-revalidate), routes onsale traffic to the Waiting Room before it can reach the gateway, and runs the bot-manager scoring (device fingerprint + IP-ASN reputation + behavioral signals). Drops 80–98% of bot traffic at the edge before it consumes origin capacity.

Why it exists. Need a tier that absorbs read load AND blocks bots before the inventory tier sees them. Rejected 'rate-limit at the gateway' because by then the bot has already consumed a TLS handshake, a token validation, and a service hop — Queue-it's published numbers show one onsale blocked 2M+ bots (98% of traffic); at gateway-tier cost that's 2M wasted requests through 4 hops.

When it fails. Fastly PoP outage (the 2021-06-08 49-min global outage is the cited precedent). Origin sees a 20–50× spike as the edge cache evaporates. Mitigation: DNS failover to Cloudflare CDN with a 5-min TTL on the public name (the load-bearing DR knob — anycast PoP recovery is not a failover we control); origin sized for ≥3× steady-state for ≥60 min.

Service · Waiting RoomAWS Lambda + API Gateway (deliberately isolated from main fleet)

Holds users in a fair queue when demand >> capacity. Pre-onsale: enrolls users into the room. T-0 (onsale): randomly shuffles the queue (fairness-by-lottery, not arrival time) and starts admitting at the rate the inventory tier can absorb. Mints a signed admission token (HMAC, 5-min TTL) for each admitted user and stores it in the Token Store. The token is required by every subsequent call into the inventory tier.

Why it exists. Need to throttle 3.5M concurrent users into ~100 concurrent shoppers per inventory unit. Rejected 'rate-limit at the gateway' because rate-limit returns 503 with no fairness — users mash refresh, the system becomes a retry storm, and arrival-time bias gives bots the win every time. Ticketmaster SmartQueue's published 13B-bots-blocked-across-17K-events stat is the cost of NOT having this tier. Following SeatGeek's published pattern: this is a deliberately-isolated AWS Lambda + DynamoDB + Fastly stack that does not share fate with the inventory tier — the thing protecting the thing on fire cannot share fate with the thing on fire.

When it fails. Waiting Room own outage: cascades into the main fleet because traffic floods to the gateway. Mitigation: stack isolation (SeatGeek's published lesson — separate AWS account, separate VPC); the inventory tier's emergency-shed mode at the LB returns 503 with Retry-After and a 'currently in maintenance' page. Detected by wr.lambda.5xx > 1% for 60s and gw.unauthenticated_rps > 5x baseline.

KV Store · Waiting Room TokensDynamoDB (on-demand) with TTL attribute

Two item types in one table. (1) token items: active admission tokens minted by the Waiting Room — keyed by token_id, attributes (event_id, user_id, issued_at, expires_at, used_steps[]). Hot path: API Gateway GETs the token on every authenticated call. DynamoDB's native TTL attribute auto-expires tokens — no janitor job. (2) rate-limit items: sliding-window counters keyed by (user_id, window_minute) and (ip_asn, window_minute) — the gateway atomically increments and reads back the count to enforce the 60-holds/min and 6-confirms/min global ceilings (24 gateway replicas cannot enforce a global ceiling from local memory).

Why it exists. Need a globally-replicated, low-latency token store that survives a region failure. Rejected Redis because Redis Cluster has 5s RPO on async failover (config'd as redis-cluster preset) — losing 5s of recently-minted tokens during a regional cutover means the queue admits N users who can't reach the inventory tier. DynamoDB Global Tables replicates asynchronously across 3 regions (multi-active, last-write-wins, typically sub-second lag — there is no synchronous multi-region write mode). We pick it anyway because that sub-second replication RPO is far tighter than Redis's 5s failover window, LWW safely resolves concurrent multi-region token writes, and the admission token is idempotent — so a token lost in the sub-second window during a regional cutover just makes that user re-request (no double-admit).

When it fails. DynamoDB regional outage (us-east-1 2021-12-07). Other regions stay live; admission decisions stall ~5s while clients fail over via Route 53 health checks. Critical: do NOT route admission to a different region without invalidating in-flight tokens — that's how SeatGeek's published runbook covers it.

API GatewayEnvoy + custom auth filter

Validates the waiting-room admission token (HMAC + DynamoDB lookup) on every authenticated call, applies per-user / per-IP rate limits, terminates TLS, and routes to the right backend (search-svc, hold-svc, confirm-svc). Strips raw client headers, injects internal correlation IDs, and enforces a global X-Idempotency-Key requirement on every POST to /confirm.

Why it exists. Need ONE place to enforce token validation, rate-limit, and idempotency-key presence across 4+ backends. Rejected 'each service auths itself' because every backend would re-implement token validation; drift would let bots through on the service that updated late. Envoy's per-route rate-limit + ext_authz integration is the published pattern; idempotency-key gating happens here so a missing key never reaches the confirm service.

When it fails. Bad deploy to Envoy config (the Cloudflare 2022 retro pattern). Mitigation: progressive rollout (1 → 5 → 25 → 100% over 30 min), automatic rollback on 5xx rate > 0.5% for 60s. Detected by gw.5xx_rate and gw.request_duration_p99 against the dual SLO window.

Service · SearchJava + Spring + OpenSearch client

Handles GET /search, GET /hotels/:city/:date, and GET /events?q=.... Reads from Search Index (Elasticsearch), applies business ranking (boost active inventory, demote sold-out), runs A/B tests, and returns paginated JSON ~20–50KB per page. Stateless. Reads (not writes) availability hints from the Hold Store on the detail-page enrichment path.

Why it exists. Need a ranking + business-logic layer that's NOT in the inventory tier's path. Rejected 'query Elasticsearch directly from the gateway' because ES queries lack our business ranking (active inventory boost, sponsored placement, seasonal promotions), A/B test machinery, and result deduplication across overlapping searches; pushing that into ES script-fields gives you 10× the latency on the same node count.

When it fails. Bad deploy ranking model returns garbage results. Detected by search.zero_result_rate > 5% and search.click_through_rate.delta < -20%. Mitigation: feature-flagged ranking, instant rollback via flag flip — does NOT require a redeploy.

SearchElasticsearch 8 (managed: Elastic Cloud)

Inverted index over hotels, events, properties, dates. Faceted search (city + date-range + price + amenities), geo-search, full-text. Refreshed every 1s from the inventory primary via CDC → Kafka → indexer. Returns ranked document IDs + denormalized display fields (name, price, image URL, available-flag-as-of-1s-ago).

Why it exists. Need full-text + multi-dimensional faceted search at 100K QPS. Rejected 'search from Postgres' because GIN indexes + full-text + geo + range filters on a single Postgres query at 50–100K QPS is infeasible at our row count (~100M active properties × 365 dates = 36B unique key combinations); specialized inverted index gives sub-100ms p99 where Postgres gives 5–30s. Following Skyscanner's published pattern: ES as a data service for hotel offers.

When it fails. Cluster red status — one shard's primary + all replicas down (the Stripe 2019-07 gray-failure shape). Detected by es.cluster_status != green for 30s and es.indexing_lag_seconds > 30. Mitigation: snapshot-restore from S3-backed snapshot; meanwhile Search Service falls back to a small in-process cache of last-known-popular results (degraded UX, NOT a hard failure).

Service · HoldGo + Redis client (redis/go-redis)

Owns the hold lifecycle. POST /holds runs a Lua script in the Hold Store that atomically: (1) checks the seat/room/date isn't currently held; (2) writes the hold record keyed by hold:event:E:seat:S (or hold:property:P:date:D) with a 5–15 min PX TTL and the user's holder-token; (3) writes a pointer key hold:id:<hold_id> → the seat key with the same PX TTL so release/confirm can locate the hold from the hold_id; (4) returns the hold ID + holder-token. DELETE /holds/:id runs a Lua check-and-delete that verifies the holder-token before removing. GET /availability/:event/:seat reads the same key. Subscribes to Kafka bookings.confirmed and DELs the confirmed hold's key (holder-token-checked Lua) — the outbox-driven release; the PX TTL is the backstop.

Why it exists. Need a policy + observability layer in front of Redis. Rejected 'client writes directly to Redis' because Redis Lua scripts can't run per-user quota checks, can't emit structured hold.placed audit events for the abuse pipeline, and can't enforce the rate-limit semantics consistently across web/mobile/3rd-party — Hold Service runs those policies and emits the event, Redis runs the lock.

When it fails. Hot-event hot-shard meltdown — one event has 10–17K hold-attempts/sec and ALL of them write to the Hold Store. Detected by redis.shard.write_qps.stddev_ratio > 8x (the hot shard climbs while peers idle). Mitigation: pre-shard the event into seat-block sub-shards (Section 105 rows 1–10 on shard A, 11–20 on shard B); the hot-event escape valve re-routes to a dedicated event-local Redis cluster when the predicted heavy-hitter sketch fires.

Cache · HoldsRedis 7 (cluster mode)

Authoritative store for active holds — Redis keys with TTL ARE the inventory lock. Per-seat / per-(property, date) keys, value carries the holder-token and metadata. Hot path: 50–100K SET NX PX ops/sec at peak. Built-in PX expiry means abandoned holds free themselves with NO janitor job. AOF persistence ON (everysec) so a single-node crash loses at most 1s of holds — and HLL/Redis hold-token semantics make a lost hold a free hold (idempotent: re-creating the same hold is fine).

Why it exists. Need a contended atomic write at 50–100K QPS with built-in TTL. Rejected 'store holds in the inventory DB' because Postgres single-row INSERT ceiling is ~2K/s on strong hardware (research: capacity § 'row-level pessimistic locking breaks at ~1–2K writes/sec') — 50–100K holds/sec would melt the WAL fsync ceiling. Rejected DynamoDB (slower per-op p99 at this throughput, expensive at this volume — ~$5K/day at peak vs ~$200/day for Redis). Redis cluster's per-slot atomic Lua + native PX expiry is what makes hold-then-confirm tractable.

When it fails. Primary shard fails — RPO=1s of recent holds lost. Idempotent semantics make re-creating those holds safe: client retries and the SET NX succeeds because the lost key is gone. The user-visible hand-wave: a user who held seat 14C on the failed primary 1s before crash reads hold_not_found from the promoted replica and may race another user for the same seat during the 1s gap. The booking-row UNIQUE constraint catches the duplicate confirmed booking (defense in depth) but the UX is 'I held that, now someone else got it.' Mitigation we accepted (vs. sync replication at 2-3ms/write cost): bounded blast radius (1 shard ÷ 16 = 6% of holds during the gap), client-side 'hold lost — pick a new seat' UX prompt. Critical: a noeviction OOM IS a hard outage — redis.used_memory_pct > 85 pages immediately because new holds 503 (vs. evicting existing holds, which would silently double-book).

Service · ConfirmJava + Spring + Stripe SDK

Owns POST /confirm. The orchestrator: (1) claims Idempotency-Key in the Inventory DB's idempotency_keys table with a unique constraint, in its own txn (fails fast on replay → returns cached response); (2) verifies the hold still lives in the Hold Store via a Lua check; (3) calls the Fraud Service synchronously (≤200ms, fail-open on circuit-breaker — see fraud node); (4) calls the Payment Provider with the same idempotency-key forwarded; (5) on charge success, opens a SERIALIZABLE Postgres transaction that INSERTs the booking row, INSERTs the outbox row for fan-out, and UPDATEs the idempotency row with the cached response — the Redis hold release is driven by that outbox event (CDC → Kafka → hold-svc), with the PX TTL as backstop; (6) returns 201 with the booking ID. Saga compensation path (5b): if Stripe charge succeeds but the inv-db SERIALIZABLE txn fails (deadlock, primary failover at the wrong instant), publish bookings.failed_post_charge straight to Kafka with acks=all carrying {charge_id, idempotency_key, user_id} — NOT via an outbox row on the shard that just failed (the failover case would share fate with it). At-least-once is safe: Compensation Worker dedups on idempotency_key, and the PSP webhook (charge.succeeded with no matching booking after N minutes) is the durable backstop if confirm-svc dies before publishing. Compensation Worker picks it up and issues a refund — see compensation-svc.

Why it exists. Need ONE atomic boundary that ties the payment authorization to the booking commit, with idempotency on top. Rejected 'client calls Stripe directly + then calls our DB' because the user could be charged but the booking write could fail — and there's no compensating undo because the charge has already cleared Stripe's books; you'd owe the user a refund out-of-band. The dual-write divergence is the classic 'distributed transaction we don't actually want.' Following Stripe's published pattern: the confirm service owns the idempotency contract and the booking write is gated on a durable claim.

When it fails. Payment provider hangs (the cited Visa Europe 2018 retro pattern — provider returns no error, just hangs). Detected by payment_svc.outbound.latency.p99 > 5s for 3m AND payment_svc.http_client.pool_in_use_pct > 90. Mitigation: aggressive timeout (3–5s NOT the default 30s), circuit breaker on payment client, separate connection pools per provider, dual-provider fallback (Stripe → Adyen) via the abstraction layer.

SQL DB · Inventory + IdempotencyPostgres 16 (Patroni + Citus sharding, AWS RDS or self-managed)

System of record. Three tables in one shard: bookings (the durable booking row), outbox (events to publish post-commit), idempotency_keys (client-supplied key + cached response, unique constraint). The idempotency claim runs in its own txn BEFORE the charge (it gates the flow); the post-charge SERIALIZABLE txn INSERTs booking + outbox and UPDATEs the idempotency row with the cached response; the outbox is published asynchronously by CDC Relay. Sharded by event_id (ticketing) / property_id (hotels) so a hot event lives on one shard and doesn't bleed to peers.

Why it exists. Need a system-of-record that supports (a) ACID across booking + idempotency + outbox in one txn, (b) horizontal sharding by event/property so hot events don't melt the global write tier, (c) standard SQL escape hatches for ops (top 100 fraud-flagged bookings, schema migrations, point-in-time restore). Rejected Cassandra LWT (4-round-trip Paxos per write = ~10× latency at this contention; published Monzo 2019 retro shows the operational cost). Rejected Spanner (cross-shard txn is fine but +50% per-write latency and 3× cost; we don't need geo-replicated writes). Following Booking.com's published massively-horizontal sharded MySQL pattern and Brandur's published Postgres-based idempotency-key implementation.

When it fails. Single-shard primary failure (the Stripe 2019-07 retro shape — gray failure cascades across services sharing the shard). Detected by pg.shard.up != 1 and pg.replication_lag_p99 > 2s. Mitigation: Patroni auto-failover to in-AZ sync follower — AZ-RTO = 30s (that 30s is a write outage for 1/16 of events); idempotency keys make retries safe post-promotion. Regional-RTO = 10 min, operator-promoted to the cross-region DR follower (the GitHub 2018-10-21 split-brain MySQL retro is why we do NOT auto-promote on region failure). Note: a us-east-1 region outage drops the inventory primary entirely; this canonical scopes inventory to one regulatory geography — for GDPR / India-DPDP a separate inv-db-eu would be needed, with a gw routing rule on property residency.

External · Payment ProviderStripe + Adyen (dual-provider abstraction)

Authorizes + captures the charge. Confirm Service calls POST /charges with Idempotency-Key: <client-uuid>; Stripe persists the first response (status + body) keyed by it and replays it for 24h (Stripe's published contract). Returns charge_id on success; we attach it to the booking row.

Why it exists. We don't move money ourselves — PCI scope, settlement, and fraud-scoring belong with the PSP. Rejected building our own payment rail because PCI Level 1 compliance is a 6–12 month audit cycle per acquirer + per card scheme + recurring penetration tests; using Stripe drops us to PCI SAQ-A scope. Dual-provider (Stripe primary, Adyen warm fallback) is the cited mitigation for the Visa Europe 2018-class single-provider outage.

When it fails. Provider hard outage (Stripe down). Detected by payment_svc.provider=stripe.error_rate > 50% for 60s. Mitigation: automatic failover to Adyen; holds keep being created (Hold Store still works); confirms route to Adyen — published UX: 'Pay with backup processor' banner. If BOTH providers are down: auto-release-hold mode at 60s lock-up so inventory recycles instead of locking up under un-confirmable holds.

Worker · CDC RelayDebezium + Kafka Connect

Tails the Postgres logical-replication slot on each Inventory DB shard. Reads the outbox table rows (and only the outbox — filtered via Debezium's outbox-event-router SMT), decodes them, publishes to the right Kafka topic (bookings.confirmed, holds.released, etc.), and commits the Kafka offset before advancing the WAL slot. One Debezium connector task per shard (16 shards → 16 tasks spread across the 6 Kafka Connect workers; Connect rebalances tasks to surviving workers on failure).

Why it exists. Need a way to atomically tie the booking commit to the downstream Kafka publish. Rejected 'confirm-svc dual-writes to Postgres AND Kafka' because dual-write is NOT atomic — the DB commits and then the Kafka publish fails (or vice-versa), and you have permanent divergence with no known reconciliation. The transactional-outbox pattern (Confluent's published course) is the only known-correct way: write to the outbox in the SAME transaction as the booking, then a separate CDC relay publishes from the durable log.

When it fails. Relay stalls (consumer crash, disk pressure, Kafka network blip). WAL slot grows; disk pressure on the inventory primary threatens write availability. Detected by pg.replication_slot_lag_bytes > 10GB (P1) and oldest_unpublished_outbox_age_seconds > 30 (P2). Mitigation: scale relays horizontally per shard; if catastrophic, swap to a fresh slot + replay from WAL archive offset (Debezium's published recovery playbook).

Stream · Event BusKafka 3.7 (Confluent / MSK)

Durable event bus for everything that happens post-booking. Topics: bookings.confirmed (P0 priority, partitioned by booking_id), bookings.canceled, holds.released, inventory.changed (for search-index refresh), partner.sync (per-partner priority lanes). 7-day retention enables replay after a consumer bug or schema mistake.

Why it exists. Need a durable, replayable event log between the inventory write and N consumers. Rejected RabbitMQ because once a message is acknowledged it's gone — reprocessing after a consumer bug requires reproducing events from the source-of-truth, which leaks the outbox layer's responsibility back into application code. Kafka's 7-day retention + offset semantics let us replay a downstream consumer's window without any application changes.

When it fails. Single broker failure — automatic leader re-election within 5s (Kafka's published RTO). Detected by kafka.under_replicated_partitions > 0 AND bookings.confirmed.publish_age_p99 > 30s (the latter is the business-pain pager: 'user paid but no confirmation email for 30 min'). Mitigation: rolling restart preserves cluster availability; sustained 5min broker outage triggers add-replica via Cruise Control rebalance.

Worker · Fan-outGo + Kafka consumer group + per-priority pool

Consumes from Kafka topics by priority. P0 (notifications.p0.ticket-delivery): emits email + push notifications via SES / SendGrid + FCM/APNs — the user gets their ticket / booking PDF. P1 (inventory.changed): writes the seat-sold/room-booked flag back to Elasticsearch via bulk indexer so search reflects sold-out within ~1s. P3 (partner.p3.marriott-pms): pushes to hotel PMS / barcode-mint service / CRM. Priority isolation via separate consumer groups + thread pools.

Why it exists. Need to do side effects (email, push, partner sync) AFTER the booking commits, WITHOUT blocking the confirm response. Rejected 'sync fan-out from confirm-svc' because a SES/SendGrid timeout would block the confirm response — user sees 'your card was charged but no confirmation', which is the catastrophic UX failure mode of the dual-write divergence. Async fan-out via Kafka unblocks the confirm and gives each downstream consumer its own retry budget.

When it fails. P3 partner consumer hits a poison message (malformed booking event) and stalls the partition. Detected by consumer.same_offset_for_seconds > 60. Mitigation: bounded retries (3), then route to DLQ; per-partition isolation means a poisoned partition doesn't block siblings; schema validation at the outbox writer side (NOT just consumer) catches most malformed events before they hit the bus.

Service · Fraud ScoringJava + Sift SDK + Stripe Radar fallback

Called synchronously by Confirm Service before the Stripe charge. Sends {user_id, booking_value, ip, device_fingerprint, velocity_signals} to Sift / Stripe Radar; returns a 3-tier verdict (accept / review / block). Block → confirm-svc returns 402 fraud_decline (industry-standard ~3–5% of legitimate confirms hit this path). Review → confirm-svc still charges but flags the booking for ops review. Accept → confirm-svc proceeds to charge.

Why it exists. Need a pre-charge fraud check at 10K confirms/sec — chargebacks have a 60-day settlement window and a ~1–3% chargeback rate at our class of business eats the margin. Rejected 'rely on Stripe Radar alone' because Radar fires AFTER the charge succeeds, and a chargeback unwinds the inventory commit days later — by then the seat / room can't be resold. Rejected 'no fraud scoring' because at the 20:1 bot ratio of an onsale, scalper-friendly volumes leak through. Pre-charge scoring at <200ms is the load-bearing mitigation.

When it fails. Fraud-svc cascade failure during onsale (Sift's published 2023 outage pattern). Detected by fraud_svc.5xx > 0.5% for 60s AND fraud_svc.latency_p99 > 500ms. Mitigation: circuit-breaker opens, confirm-svc proceeds in fail-open mode; ops review queue inherits the burden. Trade-off explicitly accepted: brief widow of higher chargeback risk vs. blocking all confirms during a fraud-svc blip.

Worker · Compensation (Saga)Go + Kafka consumer + Temporal-style state machine

Subscribes to Kafka bookings.failed_post_charge (emitted by confirm-svc when the inv-db SERIALIZABLE txn fails AFTER Stripe charge succeeded). For each event: (1) calls Payment Provider POST /refunds with a refund-scoped idempotency-key (refund:{payment_intent_id}) + charge_id to reverse the charge; (2) writes a compensation_log row in inv-db with {idempotency_key, charge_id, refund_id, status}; (3) emits bookings.refunded for downstream notification. Idempotent via idempotency_key — replays are safe.

Why it exists. Confirm-svc cannot commit booking AND charge in one ACID transaction — payment is an external system. Rejected 'pretend it can never fail post-charge' because a 1-in-100K deadlock at 10K confirms/sec is one stranded charge every 10s during peak. Rejected 'manual ops reconciles every morning' because the user sees their card debited and no booking — a P1 incident-grade UX. Following Airbnb's published saga pattern: the orchestrator drives the workflow; the compensation step is a real worker, not a runbook.

When it fails. Refund itself fails (Stripe rejects, e.g. card closed). Detected by compensation.refund_5xx_rate > 5% for 5m. Mitigation: after 5 attempts route the event to a manual-ops queue; ops issues an ACH refund / store credit out-of-band. The cardinal contract: the user is MADE WHOLE within 24h regardless of automation success.

Service · PSP Webhook ReceiverGo + HMAC verifier + Postgres dedup table

Public HTTPS endpoint for asynchronous events from Stripe / Adyen (charge.succeeded, charge.refunded, charge.dispute.created, payment_intent.canceled). Verifies HMAC signature against per-PSP secret; claims (psp, event_id) in a Postgres dedup table (30-day TTL because PSPs replay for up to 30 days); publishes the canonical psp.event to Kafka; returns 200. Downstream: compensation-svc consumes for refund acks; fanout consumes for the customer-facing 'card charged' email; confirm-svc's cancel path consumes for chargeback-triggered cancellations.

Why it exists. Need a separate ingress for the PSP back-channel that does NOT share fate with our public API gateway. Rejected 'reuse gw + auth bypass' because a misconfigured auth rule on gw blocks PSP webhooks, which then triggers PSP retries (24h × N webhook subscribers) — published Stripe outage retro pattern. Dedicated webhook-svc with PSP-specific HMAC verification isolates the back-channel from public API auth changes.

When it fails. Bad HMAC (the cited Square 2023 cert-bundle incident pattern). Detected by webhook.hmac_failure_rate > 0.1% for 60s. Mitigation: per-PSP HMAC secret rotation runbook with manual T-30/14/7 day alerts; failing-closed (returning 401) is correct — PSP will retry.

Object Store · Audit SinkS3 + Glue catalog + Parquet (Iceberg table format)

Append-only durable log of every booking + payment event for 7-year PCI / financial regulatory retention. A dedicated Fanout-Worker consumer group subscribes to bookings.*, payments.*, and psp.event Kafka topics; batches records into Parquet files (1 file per minute per topic), writes to S3 with Glue/Iceberg schema management. Read path: Athena / Trino for ad-hoc audit queries.

Why it exists. PCI / SOX / state-AG audits demand a tamper-evident booking + payment log retained 7+ years. Rejected 'keep it all in Postgres' because 7 years × 100M bookings/yr × 2KB = ~1.4 TB raw on the OLTP cluster — drags every backup, every replica restore, every schema migration through cold data nobody queries. Rejected 'one row per event in DynamoDB' because Athena cannot meaningfully run aggregate audit queries against unindexed DynamoDB — Parquet on S3 is the published cheap-cold-store pattern (Snowflake / Databricks / Iceberg lineage).

When it fails. S3 regional outage (us-east-1 2017, 2021 incidents). Detected by s3.put_5xx_rate > 1% for 60s. Mitigation: cross-region replication to us-west-2 (lazy, eventual); buffer-and-replay from the Kafka 7-day retention if S3 is down for hours. Audit-sink is NEVER on the user-facing critical path — it can be silent for hours without user impact.

Stage by stage

The same 10 stages the workspace walks, answered.

01Clarifications

What would you ask before drawing a single box?

Typical clarifications to surface:

  • Ticketing or lodging? Both fit this pattern but tunables differ. Tickets: hold TTL 5–10 min, hot-event skew ~99.9% on day-of-onsale, virtual waiting room is mandatory. Hotels: hold TTL 10–15 min, hot-key skew on (property, date) ~100× the mean, search is more important.
  • Onsale vs steady-state? Onsale (concert, holiday) traffic is 100–1000× steady-state. The architecture must size for the burst without paying for it at steady-state.
  • Search expectations? Faceted multi-attribute (city + date + price + amenities) → Elasticsearch. Catalog only → a hash lookup is enough.
  • Multi-region? Most ticketing is single-region (the inventory tier is CP). Hotels span regions but inventory writes still go to a home region per property; only search is multi-region.
  • Bot tolerance? Assume 20:1 bots-to-humans at onsale; block at the edge before the queue.
  • Payment provider? Stripe / Adyen / Braintree are the standard external sync calls. Single-provider = single point of failure; dual-provider is the load-bearing mitigation against the Visa Europe 2018-class outage.
  • Hold TTL semantics? Is the seat immediately available again on hold expiry, or is there a brief grace period? Choose 0s grace (immediate) — the simplicity dominates the rare "I was about to pay" UX.

Assumptions to state:

  • Peak concurrent users in queue (Taylor-class): 3.5M.
  • Tickets per hot event: 50K, sell-through window 5 min.
  • Oversubscription (attempts per success): 50–100×.
  • Hold TTL: 7 min default.
  • Confirm conversion: 60% of holds become bookings.
  • Fan-out: 6–8 side effects per confirmed booking (email, push, PMS, CRM, analytics, fraud-scoring, indexing).

02Functional reqs

What must this system actually do?

  • Search and browse inventory (faceted: location + date + price + amenities).
  • Show inventory detail with a live availability check (the cached search result might be stale; the detail page is authoritative-ish).
  • Place a soft hold on a seat / room / (property, date) for N minutes; cancel a hold; auto-release on TTL expiry.
  • Confirm a held inventory unit via payment; receive a booking confirmation. Must be idempotent.
  • Cancel a confirmed booking (within the property's / event's policy window); release the inventory and refund.
  • Deliver the booking artifact (ticket PDF + barcode for events; reservation email + PMS push for hotels).
  • For hot-event onsale: hold users in a virtual waiting room until inventory tier capacity is available; admit them in a fair (lottery-shuffled, not arrival-order) sequence.

03Non-functional

What must it promise about speed, uptime and correctness?

  • Availability: 99.95% on the read path (search + detail). 99.9% on the confirm path during onsale (write tier).
  • Latency: Search p99 200ms; availability p99 100ms; hold p99 300ms; confirm p99 5–10s (payment-provider-dominated); waiting-room admission p99 50ms.
  • Durability: Bookings forever — 7+ years for financial / PCI retention. RPO=0 on bookings (we run synchronous replication on inventory primary). RPO ~1s on holds (acceptable: a lost hold is a free hold; idempotency makes retry safe).
  • Consistency: Search is eventually-consistent (1s ES refresh). Hold-then-confirm is linearizable on the hold key (Redis-Cluster per-slot atomic Lua). Booking commit + idempotency claim + outbox publish is one SERIALIZABLE transaction.
  • Scalability: Horizontal on every tier. Inventory primary sharded by event_id / property_id so hot events don't melt the global write tier.
  • Security: PCI scope is SAQ-A (we never see raw card data — only Stripe tokens). Bot blocking at edge. Per-identity hold quotas. Idempotency keys client-minted to defend against retry-induced double-charges.

04Capacity estimation

How much load and data does this have to hold?

Anchor metric: Taylor Swift Eras Tour onsale, Nov 15 2022 — 3.5M Verified-Fan pre-registrations, 14M users showed up for a system sized for 1.5M, 3.5B total requests during the onsale window, ~4× Ticketmaster's prior peak.

Per-tier QPS (formulas, then numbers — assumptions: 50K-seat stadium, 5-min sell-through, 60× oversubscription):

  • Search (sustained / peak): DAU × pageviews / 86,400 × peak_factor. 10M DAU × 10 PV/day × 4 ≈ 5K QPS sustained, 50–100K peak. Booking.com globally is ~1M+ search QPS at peak. Latency: p50 50ms, p99 200ms.
  • Availability check: ~5–10% of search → 1K sustained, 5–10K peak. p99 100ms.
  • Hold (per hot event): ticketsPerEvent / sellThroughSeconds × oversubscriptionFactor = 50,000 / 300 × 60 ≈ 10–17K hold-attempts/sec on ONE event. Aggregate across simultaneous onsales: 50–100K hold QPS. p99 300ms.
  • Confirm: successfulHoldQps × confirmConversion — successful holds are bounded by sell-through, not attempts: per hot event 50,000 / 300 × 0.6 = 100 confirms/sec; aggregate across simultaneous onsales + steady-state lodging ≈ 1–1.7K successful holds/sec → ~0.6–1K organic confirms/sec. We size the confirm tier at a ~10K confirms/sec design peak (~10× headroom — a mis-predicted onsale calendar must not brown out the write tier); every downstream sizing figure uses the 10K design peak. p99 5–10s (Stripe-dominated).
  • Virtual waiting room: ~tens-of-K admissions/sec into single events. p99 50ms.

Write-path bandwidth at each service tier (the bytes-economics check):

  • Hold Service: 50K QPS × 3KB req/resp = 150 MB/s through 32 replicas → 5 MB/s/replica. Fine.
  • Confirm Service: 10K confirms/sec × 5KB = 50 MB/s through 32 replicas → 1.6 MB/s/replica. Fine.
  • Inventory DB: 10K commits/sec × 2KB row × 16 shards = 1.25 MB/s/shard WAL write. Fine.
  • Payment provider call: 10K confirms/sec × 22s timeout window = 220K concurrent requests × 2KB = ~440 MB in-flight but mostly idle bytes; Stripe absorbs.

Bytes economics — explicitly checked: no edge in this canonical carries user-supplied bodies > 100 MB. All payloads are JSON < 10KB. Service tiers correctly sit in the data path because there is no byte economy to optimize. This is NOT a YouTube / image-upload problem; do not import direct-to-storage patterns from those canonicals.

Storage growth:

  • Bookings (7yr retention): 100M bookings/yr × 2KB × 7 = 1.4 TB raw, ~4 TB with indexes. Trivially fits across 16 sharded Postgres primaries.
  • Holds (TTL'd, working set): 1M concurrent holds × 500B = 500 MB working set in Redis. One Redis Cluster of 16 shards × 4 GB each = 64 GB total memory — plenty of headroom.
  • Idempotency keys (24h TTL): 20K QPS × 86,400 × 200B = ~350 GB. Across 16 Postgres shards = 22 GB/shard. Fine.
  • Audit log + payment events (7yr): ~5–10× booking volume → S3 + Parquet.

Fan-out per confirmed booking: 6–8 side effects → at 10K confirms/sec peak that's 60–80K async events/sec through Kafka. Trivial at RF=3, 32 partitions, 3 brokers (cluster headroom is published at 100 MB/s/broker; 60–80K × 1KB = 60–80 MB/s aggregate).

Architectural flip points (where the design changes shape, not just numbers):

  • Waiting Room becomes mandatory at ~50K concurrent users in queue. Below that, a CDN + rate-limit suffices. Taylor-class (3.5M queued) needs a dedicated, isolated tier.
  • Row-level pessimistic locking breaks at ~1–2K writes/sec per hot row. Postgres single-row INSERT ceiling is ~2K simple-insert TPS on strong hardware. Above that move to pre-sharded seat-blocks (one row per seat-block, parallel locks) OR Redis-Cluster per-slot atomic Lua OR token-based admission (pre-mint N reservation tokens per inventory unit). We use Redis-Cluster Lua for holds (50–100K QPS achievable) and sharded Postgres for the durable commit (10K commits/sec achievable).
  • Single-region inventory becomes unviable when (a) regulation forces residency (GDPR / India DPDP) or (b) cross-region p99 > 150ms makes the hold latency budget infeasible. Then partition inventory by region-of-origin with no cross-region locks.

05API design

What does the outside world call, and what comes back?

GET /search?city=NYC&checkin=2026-06-01&checkout=2026-06-05&adults=2
200 {
  "results": [
    { "property_id": "p_42", "name": "Hilton Times Square", "price_cents": 39500,
      "available_hint": true, "available_as_of": "2026-05-11T15:42:01Z" },
    ...
  ],
  "page": 1, "total": 248
}
GET /availability/:event/:seat        # tickets
GET /availability/:property/:date     # hotels
200 { "available": true, "hold_id": null, "ttl_remaining_seconds": null }
200 { "available": false, "currently_held_until": "2026-05-11T15:49:01Z" }   # held by someone else — a read, not a conflict
POST /holds
Content-Type: application/json
Authorization: Bearer <waiting_room_token>
X-User-Id: u_993
X-Hold-Intent-Key: 5e3a8b1c-9f44-4e2a-b8d3-c1f2a3b4c5d6   # client-minted UUID — gates SETNX so retries return same hold_id

{ "event_id": "e_taylor_eras_atl_2026", "seat_id": "sec105_row12_seat14" }

201 Created
{ "hold_id": "h_abc123", "holder_token": "0f1e2d3c-4b5a-6978-8a9b-c0d1e2f3a4b5",
  "expires_at": "2026-05-11T15:49:01Z", "ttl_seconds": 420 }
# holder_token is the release/confirm verifier stored in the Redis value — without it the
# Lua check-and-delete has nothing to check against (stolen-lock defense)

409 Conflict      # already held by someone else
{ "error": "seat_held", "available_again_at": "2026-05-11T15:49:01Z" }
POST /confirm
Content-Type: application/json
Authorization: Bearer <waiting_room_token>
X-Idempotency-Key: 73c8e8e0-3a2d-4f5b-8d2e-1c2b3a4d5e6f   # client-minted UUID per Pay intent

{
  "hold_id": "h_abc123",
  "payment_token": "stripe_tok_visa_xxxx",
  "billing_zip": "10001"
}

201 Created
{ "booking_id": "b_xyz789", "charge_id": "ch_stripe_xxx",
  "ticket_pdf_url": "https://...", "delivered_at": null }

409 Conflict
{ "error": "hold_expired" }       # the hold TTL'd while the user was paying

504 Gateway Timeout + Retry-After  # confirm-svc contract: client retries on 504 with the same idempotency key → cached 201 within 24h
POST /holds/:hold_id/release      # explicit cancel — also: auto on TTL
{ "holder_token": "0f1e2d3c-4b5a-6978-8a9b-c0d1e2f3a4b5" }   # Lua check-and-delete verifies this before DEL

200 { "released": true }
POST /bookings/:booking_id/cancel     # ownership-checked against the authenticated user
X-Idempotency-Key: 8a2f4c6e-1b3d-4e5f-9a7b-2c4d6e8f0a1b   # client-minted UUID per cancel intent

200 { "booking_id": "b_xyz789", "status": "canceled", "refund_status": "pending" }
     # one txn: status=canceled + outbox row — the refund and the inventory release
     # ride the existing outbox → Kafka (bookings.canceled) path

409 Conflict
{ "error": "outside_cancellation_window", "cancellable_until": "2026-05-25T00:00:00Z" }

06Data model

What gets stored, and what is it looked up by?

holds (Redis, TTL'd):

key:    hold:event:<event_id>:seat:<seat_id>
value:  {"hold_id":"h_abc123","user_id":"u_993","token":"<uuid>","placed_at":"..."}
PX:     420000                      # 7 minutes

key:    hold:id:<hold_id>           # pointer — release/confirm locate the seat key from the hold_id
value:  hold:event:<event_id>:seat:<seat_id>
PX:     420000                      # same TTL as the seat key

bookings (Postgres, sharded by event_id / property_id):

fieldtypenotes
booking_iduuidPK
event_or_prop_idbigintshard key
seat_or_date_idtextunique within event_or_prop_id
user_idbigint
charge_idtextfrom Payment Provider
statusenumconfirmed / canceled / refunded
created_attimestamp
amount_centsbigint

UNIQUE INDEX (event_or_prop_id, seat_or_date_id) WHERE status = 'confirmed' — the load-bearing oversell defense: even if every other layer fails, the DB will reject a duplicate confirmed booking.

idempotency_keys (Postgres, same shard as booking):

fieldtypenotes
keyuuidclient-minted UUIDv7
key_datedatefrom the key's embedded timestamp; partition key
user_idbigint
request_hashbyteahash of request body — mismatch returns 409
response_codeintthe cached status code
response_bodyjsonbthe cached response
created_attimestamp
expires_attimestampcreated_at + 24h

PK is (key, key_date) — Postgres requires the partition column in every unique constraint on a declaratively partitioned table. key_date is a pure function of the UUIDv7 key, so a retried key always targets the same daily partition and concurrent INSERTs still serialize on the key. ON CONFLICT DO NOTHING RETURNING id is the gate.

outbox (Postgres, same shard, same txn):

fieldtype
outbox_idbigint
topictext
payloadjsonb
created_attimestamp

The CDC Relay reads from here and publishes to Kafka, then advances the WAL slot offset.

Database choice — recommended: Sharded Postgres for the system of record. Why: ACID across booking + idempotency + outbox in ONE transaction; mature operational tooling; standard SQL escape hatches for ops queries; horizontal sharding by event_id / property_id keeps hot events local to one shard. Following Booking.com's published massively-horizontal sharded MySQL pattern, adapted to Postgres for SERIALIZABLE isolation.

Why not Cassandra LWT? Cited Monzo 2019-07-29 retro: light-weight transactions are 4-round-trip Paxos and the operational cost at this contention is severe (auto_bootstrap=false on scale-up produced empty quorum reads → card-payment outage). Postgres single-shard SERIALIZABLE on ~625 writes/sec/shard is the simpler, cheaper, safer choice.

Why not Spanner / CockroachDB? Fine choice if you need geo-distributed serializable writes, but the cost is +50% per-write latency and 3× the infrastructure cost. We don't need cross-shard atomicity (booking is single-shard by construction); the cost isn't worth it.

07High-level design

Which components handle a request, and in what order?

Architecture summary (the load-bearing flow):

  1. CDN + Bot Edge (Fastly + Cloudflare Bot Mgmt): absorbs 70–98% of read traffic, blocks bots before origin. Routes onsale traffic to the Waiting Room before the inventory tier sees it.
  2. Waiting Room (separate AWS Lambda stack + DynamoDB Token Store): holds users in a fair shuffled queue when demand >> capacity. Mints HMAC admission tokens. Deliberately isolated from the main fleet (SeatGeek's published pattern — the thing protecting the thing on fire cannot share fate with the thing on fire).
  3. API Gateway: validates admission tokens, enforces idempotency-key on POST /confirm, applies per-user rate limits, routes to the right backend.
  4. Search Service → Elasticsearch: read-only, eventually-consistent, multi-region. Refreshed every 1s from the inventory primary via CDC.
  5. Hold Service → Redis Cluster (Hold Store): the contended write tier. SET NX PX + Lua makes the hold atomic per Redis slot; built-in PX expiry auto-releases abandoned holds. Per-seat keys spread naturally across slots, defusing the hot-event hot-slot.
  6. Confirm Service → Inventory DB (Postgres, sharded) + Payment Provider: the orchestrator. Claims the idempotency key (unique constraint, own txn), verifies the hold lives, calls Stripe with the same idempotency key forwarded, then commits booking + outbox + cached idempotency response in ONE serializable transaction; the Redis hold release follows asynchronously via the outbox event (CDC → Kafka → hold-svc DEL), TTL as backstop.
  7. CDC Relay (Debezium) → Kafka: tails the Postgres WAL, publishes outbox rows to Kafka — the only known-correct way to atomically tie a DB commit to a downstream publish.
  8. Fan-out Worker: consumes Kafka by priority lane — P0 (ticket delivery + email), P1 (search-index refresh), P3 (partner sync / PMS / CRM). Side effects never block the confirm response.

Data flow on read (search → detail): Client → CDN → (miss) → API Gateway → Search Service → Elasticsearch → respond. CDN caches search results for 60s with stale-while-revalidate.

Data flow on hold: Client → CDN → (post-admit) → API Gateway → Hold Service → Hold Store (SET NX PX + Lua atomically reserves the seat) → respond with hold_id + TTL.

Data flow on confirm: Client → CDN → API Gateway → Confirm Service. Confirm Service: (a) idempotency claim in Postgres (ON CONFLICT DO NOTHING); (b) verify hold lives in Hold Store; (c) call Payment Provider with same idempotency key forwarded; (d) one SERIALIZABLE Postgres txn writes booking + idempotency + outbox row (single store — atomic); (e) return 201; (f) the Redis hold release is driven by that outbox event — CDC → hold-svc DELs the key (Lua, holder-token-checked), async. It is not a synchronous DELETE inside the Postgres txn (a Postgres txn can't atomically delete a Redis key; coupling them is the cross-store dual-write deep-dive #4 rejects). If the release is delayed or fails, the hold's 420s TTL reaps it and the bookings UNIQUE constraint blocks any double-confirm meanwhile. Async: CDC Relay → Kafka → Fan-out Worker → email/PMS/search-index refresh.

Multi-region posture: Single-region inventory primary (the CP tier). Multi-region read replicas for search (Elasticsearch active-active reads). Token Store on DynamoDB Global Tables (idempotent token, LWW-safe). Hold Store single-region (rebuilt from inventory primary on regional failover; ≤5s of holds lost on the async-replication tail, which is safe because hold creation is idempotent).

08Deep dives

Which part breaks first, and what do you do about it?

1. Hot-event hot-shard defense. A Taylor Swift onsale concentrates 99.9% of write traffic on ONE event for the day. The naive approach (shard by hash(event_id)) puts all of that traffic on one shard. Three layered mitigations: (a) pre-shard the event into seat-block sub-shards (Section 105 rows 1–10 on shard A, 11–20 on shard B) so contention spreads across hot shards; (b) per-seat keys in Redis (no {event:E} hashtag pinning) so each seat is its own slot in the Redis Cluster — defuses the hot-slot meltdown the concurrent-hotel-viewers canonical separately discusses for HLL; (c) hot-event escape valve — a routing layer that detects predicted heavy-hitters (Verified-Fan registration count) and routes that event to a dedicated event-local Redis cluster with its own primary, so the rest of the system isn't degraded by the hot event's load.

2. The hold-then-confirm race. User A holds seat 14C; user B sees it as available in search (eventually-consistent index, 1s lag); B clicks "buy". Where does the race resolve? At every layer:

  • Search shows "available" — fine, search is eventual.
  • Detail page does a "live check" against Hold Store — A's hold is now visible; B sees currently_held_until: ....
  • B's POST /holds is SET NX PX — Redis Lua rejects because the key exists. B sees 409.
  • If B somehow got past (race between detail page render and POST), the confirm-svc would re-check the hold under linearizable Redis Cluster semantics AND the Postgres bookings UNIQUE constraint would reject the second confirmed booking — defense in depth.

3. Idempotency under retry storms. Client mints a UUID at "open payment sheet" (not at submit). The same key is forwarded on every retry of the same intent. Confirm-svc claims it in Postgres idempotency_keys with ON CONFLICT DO NOTHING RETURNING id. If zero rows returned, the key was already claimed; load the cached response from the same row. Crucial: do not use Redis alone for dedup — the cited failure mode is Redis OOM under unbounded TTL → silent eviction → duplicates resume. Postgres + unique constraint + bounded TTL via daily partition is the durable contract.

4. Why CDC + outbox, not dual-write. Dual-write (confirm-svc writes to Postgres AND Kafka) is not atomic. If Postgres commits and Kafka publish fails, you have a confirmed booking with no downstream notification — user paid but no email, no PMS push, no analytics. If Kafka publishes and Postgres commit fails, you have a notification for a booking that doesn't exist. Both are unrecoverable without manual reconciliation. The outbox + CDC pattern (Confluent's published course) writes the outbox row in the SAME txn as the booking; the CDC relay then publishes from the durable log. Failure is bounded: relay falls behind → grafana lag alert → operator drains, never silent divergence.

5. Payment-provider outage survival. Stripe down for 10 min: dual-provider abstraction routes confirms to Adyen automatically on provider.error_rate > 50% for 60s. Holds keep being created (the Hold Store doesn't depend on payment); only confirms degrade. Critical secondary mitigation: when BOTH providers are down, hold-svc enters auto-release mode at 60s — un-confirmable holds release back to inventory instead of locking it up under load that can never complete. User-visible UX: "Payment is currently unavailable — your seat was released back to inventory; we'll notify you when payment service resumes" rather than charges + booking-undelivered.

6. Bot armies hoarding holds. Scalpers use residential-proxy networks to bypass the waiting room and hoard holds with no intent to buy. Detection: holds_expired_unconfirmed_rate > 3× baseline plus unique_residential_proxy_ASNs.per_event > 500 plus device-fingerprint entropy collapse. Mitigation: (a) Akamai/Cloudflare Bot Management at the edge with behavioral scoring; (b) per-identity hold quota (max 2 active holds per Verified-Fan ID per event); (c) velocity limits per IP-ASN bucket (no more than 10 hold attempts/min from the same ASN cohort); (d) shorten hold TTL aggressively when bot scores rise (90s instead of 7min) so unbought inventory recycles fast.

09Trade-offs

What did this design cost, and what breaks at 10×?

What we accepted (vs. the stronger alternative):

  • RPO=1s on Hold Store (Redis async replication). Sync replication on Redis adds 2–3ms/write which is not free at 50K QPS. The 1s RPO is safe because hold creation is idempotent — a lost hold simply means the user's retry recreates it.
  • Eventually-consistent search (1s ES refresh). Stronger consistency would require synchronous index updates on every booking — published Skyscanner data shows this is infeasible at our QPS. The detail-page "live check" against Hold Store covers the user-visible UX.
  • Single-region inventory primary. Multi-region active-active writes on inventory require cross-region locks or CRDTs — both prohibitively expensive. We accept that a region failure is a 30s write outage for affected shards (idempotency-safe).
  • No auto-promote on cross-region DR. The cited GitHub 2018-10-21 split-brain MySQL retro shows the cost of auto-promote — we require operator promotion (RTO ~10 min) for cross-region failover.
  • Async fan-out blocking ticket delivery. P0 (ticket delivery email) is async via Kafka — there's a ~5s gap between booking confirmation and email arrival. Acceptable: the user gets the booking_id in the 201, and the email is a side-effect; if the email fails permanently the user can re-request from the My Bookings page.

What breaks at 10× the listed scale (hypothetical 1M confirms/sec):

  • Postgres single-shard ceiling — 16 shards × 10K commits/sec = 160K/s ceiling. 1M confirms/sec needs 100+ shards OR a move to Spanner/CockroachDB.
  • Redis Cluster node ceiling — ~100K simple writes/sec per node (a slot has no ceiling of its own; the node is the unit), so the 16-node cluster caps out near 1.6M writes/sec — below the hold-attempt rate this scale implies even with per-seat key spreading. We'd need to fan-out within a single event across multiple Redis Clusters (one per seat-block).
  • CDC Relay — Debezium throughput is ~10K events/sec/replica; 1M confirms × 7 fan-out = 7M events/sec needs 700 relays. At this scale move to a different log shipper (LinkedIn's Brooklin, or migrate the outbox to a Kafka-native log).

Trade-offs we accepted

After Stage-3 critique, these trade-offs are explicitly documented (vs. closing the gap with more components):

  • Single-region inventory primary. Scope this canonical to one regulatory geography (US-East). For EU / India / China data residency, a parallel inv-db-eu would be needed, with a gw routing rule on property.region_of_origin. The single-region claim is defensible for ticketing (events have one home venue) but is the limit for hotels.
  • Hold-store async replication, 1s RPO. Sync would cost 2–3ms/write at 50K QPS — that's 100–150ms aggregate p99 budget burned for failover-window safety. We accept the 1-shard-during-1s blast radius and rely on the bookings UNIQUE constraint as the durable defense against double-sell.
  • Fraud-svc fail-open on circuit breaker. During a fraud-svc outage we accept brief widow of higher chargeback risk vs. blocking all confirms. Stripe Radar runs post-charge as a backstop.
  • No auto-promote on cross-region DR. Regional-RTO is 10 min (operator-promoted); we accept this vs. the GitHub 2018-10-21 split-brain risk of auto-promote.
  • P3 partner-sync 1h lag budget. Some hotel PMS integrations can be hours behind; user-facing confirmation email is P0 and never blocked on PMS sync.

Trace catalogue

The simulator authors 8 problem-specific traces. Each is a journey a senior engineer expects to discuss when reviewing this design:

  • hotel-booking:search-browse-cdn-hit — the 90% read path. CDN edge serves the cached search results; origin sees nothing. Budget 30ms.
  • hotel-booking:availability-live-check — detail-page "is it still available" call. Client → CDN miss → Gateway → Hold Service → Hold Store (read). The race-defeater between search staleness and the confirm path. Budget 150ms.
  • hotel-booking:waiting-room-admit — control-plane onsale flow. Client → CDN → Waiting Room → Token Store. The lottery + admission token mint. Budget 200ms.
  • hotel-booking:hold-place-redis-lua — the contended write path. Client → CDN → Gateway (token validation) → Hold Service → Hold Store (SET NX PX + Lua atomic reserve). With allowRevisit so the engine can model the idempotent retry-with-Hold-Intent-Key. Budget 400ms.
  • hotel-booking:hold-release-explicit — explicit cancel path. Client → CDN → Gateway → Hold Service → Hold Store (Lua check-and-delete; verifies holder-token before DEL). Distinct from auto-TTL release. Budget 200ms.
  • hotel-booking:confirm-pay-commit — the user-blocking path. Multi-segment: Segment 1 fraud-check + Stripe charge; Segment 2 SERIALIZABLE booking + outbox + idempotency txn. Budget 30s (Stripe-dominated).
  • hotel-booking:booking-fan-out-cdc-outbox — async multi-segment. Segment 1 CDC tails WAL → Kafka publish; Segment 2 Kafka subscribe → Fan-out Worker → ES bulk refresh + downstream side effects. Budget 30s.
  • hotel-booking:saga-compensation-refund — the load-bearing recovery path. Confirm-svc charged but inv-db txn failed → Kafka bookings.failed_post_charge → Compensation Worker → refund via Stripe with same idempotency-key → compensation_log in inv-db. Budget 60s.

Failure scenarios we model

8 chaos scenarios, each grounded in a cited real-world precedent:

  • onsale-thundering-herd (traffic) — Eras Tour Nov 2022 pattern: 3.5M concurrent queue overruns inventory tier when gate opens.
  • taylor-swift-hot-shard-meltdown (data) — 99.9% of write QPS on one event's shard while peers idle.
  • payment-provider-stripe-503 (deps) — Visa Europe 2018 single-provider outage; cited as the dual-provider mitigation rationale.
  • inventory-db-leader-failover (data) — Stripe 2019-07-10 single-shard gray failure; 30s AZ-RTO write outage on 1/16 of events.
  • hold-store-cache-stampede (traffic) — synchronized TTL expiry on hot events; Vattani et al. VLDB 2015 pattern.
  • cdc-outbox-stall (process) — Debezium consumer lag → WAL slot grows → disk pressure on inventory primary.
  • bot-army-hold-hoarding (process) — scalpers bypass waiting room via residential-proxy networks; hold-expired-unconfirmed rate spikes.
  • idempotency-table-bloat (process) — Postgres idempotency_keys partition-drop cron fails for several days; table size > 30 GB/shard; confirm latency cliffs on unique-constraint index scan.

Primary sources

  • Ticketmaster SmartQueue (blog.ticketmaster.com)
  • SeatGeek virtual waiting room (InfoQ + AWS Architecture Blog)
  • Stripe — Designing robust APIs with idempotency
  • Booking.com — Using Riak as Events Storage
  • Airbnb — Partitioning the main DB
  • Brandur — Implementing Stripe-like idempotency keys in Postgres
  • Cockroach Labs — Technical takeaways from the Ticketmaster/Eras meltdown

Now defend it

Reading a design is not the same as being able to hold one under questioning. The workspace asks the same questions an interviewer would, and the simulator disagrees with you when the diagram does not support the claim.

Work Ticketmaster / Hotel Booking yourself

More in Transactions, Concurrency & Money

Correctness when two writers collide and money is involved: serializability, two-phase commit versus sagas, hold-then-confirm, single-writer matching, and the databases that give you external consistency.