Stock Exchange / Order Matching
Worked solution

Stock Exchange / Order Matching — a worked solution

Single-writer matching. Microsecond budgets.

Try it yourself first.

You will remember almost none of this if you read it cold. The workspace walks the same 10 stages and runs the architecture you draw through a simulator, so you find out where your version breaks before you see ours.

Open the Stock Exchange / Order Matching workspace

The problem

A stock exchange is a deterministic state machine wrapped in microsecond-budgeted I/O. The dominant pattern, popularized by LMAX and now near-universal across NYSE Pillar / Nasdaq INET / CME Globex / Eurex T7, is single-writer matching behind a sequencer that linearizes every event before it touches the book. Everything else — gateways, market data, surveillance, post-trade — exists to make that single thread fast, recoverable, and provably deterministic.

The hard parts:

  • µs budget end-to-end: NIC ingress to NIC egress p99 < 50 µs across gateway + risk + sequencer + match + MD egress. p99.99 < 150 µs. HFT customers budget tick-to-trade ≤ 10 µs round trip on their side, so any internal regression is detected within minutes.
  • Determinism: replay the journal six hours later → byte-identical book. No wall clocks, no hash-map iteration order, no allocator dependence, no JIT warmup leaks.
  • Bit-identical hot-standby: same code, same JVM, same kernel, same NIC firmware. Facebook IPO 2012 shipped a secondary that had the data but ran a slightly different code path — $10M SEC settlement, $62M investor fund. State equality is necessary but not sufficient.
  • Reg SCI / 15c3-5: pre-trade risk is in the order path, not async. Knight Capital lost $440M in 45 minutes when a feature-flag flip woke 8-year-old dead code and there was no kill switch.
  • MiFID II RTS 25: ±100 µs UTC traceability for HFT venues. Lose the GPS lock for 4 minutes and you're reportable.

The reference architecture

Reference architecture for Stock Exchange / Order Matching: 20 components — Trading Participant (OMS/EMS), Colo Fabric / FPGA Tap, FIX/OUCH Gateway, Pre-Trade Risk (Reg SCI / 15c3-5), Sequencer (Aeron Cluster, Raft), Matcher (Single-Writer, In-RAM Book), Market Data Publisher (ITCH/EOBI), MD Snapshot / Recovery (TCP), Drop-Copy Stream (Kafka), Tick Store / Surveillance (kdb+), Post-Trade Processor, Cold Archive (S3 + Glacier), Ops Control / Kill Switch (Reg SCI), Shadow Book Verifier, Synthetic Traffic Injector, Reference Data / Corp-Actions, FIX Session State (FoundationDB), PTP Grandmaster (GPS-disciplined), Clearing (DTCC / NSCC / OCC), CAT Processor (FINRA CAT, LLC) — connected by 33 flows.Trading Participant (…clientColo Fabric / FPGA TapArista 7130 (FPGA) + Ci…FIX/OUCH GatewayJava + Aeron + Solarfla…Pre-Trade Risk (Reg S…In-process risk engine …Sequencer (Aeron Clus…Aeron Cluster, NVMe jou…Matcher (Single-Write…LMAX Disruptor + Java/C…Market Data Publisher…Aeron / UDP multicast (…MD Snapshot / Recover…Java + Aeron tailDrop-Copy Stream (Kaf…Kafka 3.7 (KRaft) + ide…Tick Store / Surveill…kdb+/q tickerplant + RD…Post-Trade ProcessorJava + Aeron tail + FIX…Cold Archive (S3 + Gl…S3 + Glacier Deep Archi…Ops Control / Kill Sw…Java + gRPC + audit logShadow Book VerifierIndependent toolchain (…Synthetic Traffic Inj…Canary participant + sh…Reference Data / Corp…FDB + signed bundle dis…FIX Session State (Fo…FoundationDBPTP Grandmaster (GPS-…Meinberg LANTIME M1000 …Clearing (DTCC / NSCC…externalCAT Processor (FINRA …external
20 components, 33 flows. A dashed line is an asynchronous hop. This is the reference design, not the only one that works.

Stage by stage

The same 10 stages the workspace walks, answered.

01Clarifications

What would you ask before drawing a single box?

Clarifications to surface:

  • Asset class? Equities / options / futures — different protocols (ITCH/OUCH vs OPRA/iLink), different FO message rates (options ~6× equities), different settlement (DTCC/OCC).
  • Continuous + auctions? Open / close / volatility auctions need a separate cross algorithm with its own correctness obligations (FB IPO 2012).
  • Latency tier? "Fast" venue (Nasdaq/Pillar p50 ~25µs) vs intentional speed-bump (IEX ~350µs)?
  • Co-location? All design assumes participants are colo'd in the same metro. Non-colo flow goes through "slow lane" gateways with relaxed budgets.
  • Geographic scope? Single colo (matching) + same-metro DR (mandatory). Multi-region matching is not done — latency arbitrage forbids it.
  • Regulatory regime? US (SEC 15c3-5, Reg SCI, Reg NMS, Rule 17a-4, CAT, T+1) vs EU (MiFID II / RTS 25 / RTS 7) — drives gates, retention, reporting deadlines.

Assumptions stated:

  • US equities tier-1 venue. ~15B order events/day, 6.5h regular trading hours, 252 trading days/year.
  • Peak inbound burst 5M msgs/s during open auction. Sustained ~640K msgs/s.
  • 100% determinism on replay; 99.999% RTH availability; RPO=0; mid-session RTO ≤ 30s.

02Functional reqs

What must this system actually do?

  • Accept order entry (new, modify, cancel) over FIX 5.0 SP2 + OUCH 4.2 + iLink 3 sessions.
  • Pre-trade risk gating (Reg SCI / 15c3-5): fat-finger, max position, max notional/symbol, self-match prevention, kill switch.
  • Continuous matching against a price-level FIFO book per symbol shard.
  • Run open / close / LULD / volatility auctions with deterministic crosses.
  • Publish ITCH/EOBI market data (incremental + snapshot, A+B redundant feeds).
  • Replicate every fill via drop-copy to clearing, OMS, regulator (≤100ms p99).
  • Generate T+1 settlement files for clearing (DTCC/NSCC/OCC) and CAT reports for FINRA CAT by 08:00 T+1.
  • Halt / kill switch / throttle controls accessible to ops (audit-logged, two-person rule for venue-wide).
  • Replay any seq# range deterministically for surveillance, regulatory subpoena, or DR.

03Non-functional

What must it promise about speed, uptime and correctness?

  • Determinism: 100%. Replay produces bit-identical book. Non-negotiable.
  • Availability: 99.999% during regular trading hours (~3 min/year — NYSE/Nasdaq published target).
  • Matching latency: p50 < 30 µs, p99 < 50 µs, p99.99 < 150 µs wire-to-wire.
  • Drop-copy delivery: p99 < 100 ms (regulatory + clearing).
  • MD multicast loss budget: < 0.001% of packets, A+B feed merge gives effective 0.
  • Time accuracy: ±100 µs vs UTC (MiFID II RTS 25 hard ceiling), internal SLO ±20 µs.
  • RPO: 0. Sequencer writes commit on Raft majority before ack.
  • RTO mid-session: ≤ 30s for matcher failover; sequencer leader election ≤ 5s.
  • Storage retention: 7 years immutable (SEC 17a-4 / MiFID II RTS 25), Object-Lock + Compliance mode.
  • Ordering: strict per-shard linearization. Cross-shard ordering is undefined — this is intentional.

04Capacity estimation

How much load and data does this have to hold?

Capacity walk-through — how the design follows from the load.

Inbound peak. 15B order events/day across 23,400s active session ≈ 640K msgs/s sustained. Microbursts at the open auction hit 5M msgs/s for 50–200 ms (50–100× sustained). With 16 matching partitions, peak-per-shard = 5M / 16 ≈ 312K msgs/s/shard. LMAX Disruptor benchmarks single-thread match at 6M msgs/s; we're well inside the single-writer envelope, with ~16× headroom for a hot symbol like SPY at the open.

Sequencer commit. Aeron Cluster Raft, 3 nodes per shard. The three nodes are same-cage same-rack (with separate power feeds, separate switches, ECC-rich servers), not cross-AZ — cross-AZ at 0.5–2 ms one-way puts Raft majority commit at ≥1 ms, which would blow the entire µs budget. AZ redundancy is provided by an async warm-tail follower in a second cage (different room, different cooling, different power) that consumes the leader's stream over the cage interconnect; on primary-cage loss, that follower becomes the seed for a re-elected cluster (RTO measured in tens of seconds, not 5 s — see DR posture below). Group-commit fsync to NVMe Gen4 ≈ 6 µs same-rack; theoretical write ceiling ~2M events/s/shard. We size for 1M/s/shard sustained — 3× over peak — because the sequencer commit landing on the write tail is what bounds the entire pipeline's p99.

Matcher RAM. ~10M open orders bookwide × 200 B (struct + price-level idx + client idx) × 5 (shadow copy + replay buffer + indices) / 16 partitions ≈ 625 MB/shard hot. Round to 16 GB/shard total instance RAM, of which 1–2 GB is the live book and the rest is hugepage-mlocked headroom for replay buffers, snapshot staging, and OS slack. Disk is the journal only — the book never pages out. mlockall + pre-touched pages are mandatory; a single major fault is a microsecond-budget violation.

MD egress. Outbound = inbound × 8 (level-1 + level-2 incremental + level-2 snapshot, each on A+B) ≈ 5M × 8 × 2 = 80M msgs/s peak. At ~60 B/msg ITCH this is 38 Gbps peak equities, 75 Gbps peak options (OPRA). Each shard pushes its own multicast group; total egress capacity is provisioned at 4× peak.

Journal throughput. 640K events/s × 80 B/event ≈ 51 MB/s sustained, peaks at 400 MB/s during the open. Per trading day ≈ 80 GB/shard, 1.3 TB/day venue-wide uncompressed. NVMe Gen4 sustained write 7 GB/s easily covers this; the constraint is fsync latency p99, not bandwidth.

Retention. 1.3 TB/day × 252 days × 7 years / 3 (LZ4 compression ratio) ≈ 770 TB compliant compressed. S3 Standard for 30d hot, Glacier Deep Archive thereafter. Object-Lock Compliance mode prevents deletion within the 7y window even by root credentials — SEC 17a-4(f) requires WORM media; Object-Lock counts.

Tickdb (kdb+). ~200 TB hot HDB (90 days), 5–10 PB historical (7y), columnar + LZ4. Surveillance queries (spoofing, layering, quote-stuffing) run against the day's RDB for real-time alerting and against HDB for end-of-day pattern detection.

Stream-tier flips. A single matching thread already saturates ~5M msgs/s; adding more compute to one shard never helps — you shard. Sharding is always on (16+ shards from day one), not a "scale at" tier flip. The real flip is whether you partition by symbol-hash (default, even load) or by symbol-list (lets you pin SPY/QQQ to dedicated cores when they outpace their shard). We do the latter for top-10 names.

05API design

What does the outside world call, and what comes back?

# FIX 5.0 SP2 — NewOrderSingle (35=D)
8=FIXT.1.1|9=200|35=D|49=PARTICIPANT|56=VENUE|34=42|52=20260505-13:30:00.123456789|
11=ORD12345|55=AAPL|54=1|38=100|40=2|44=185.42|59=0|60=20260505-13:30:00.123456789|
10=234

# Response — ExecutionReport (35=8) on fill
8=FIXT.1.1|9=240|35=8|...|39=2|150=F|14=100|31=185.42|17=EXEC9876|...|10=178

# FIX 5.0 SP2 — OrderCancelRequest (35=F)
8=FIXT.1.1|9=160|35=F|49=PARTICIPANT|56=VENUE|34=43|52=20260505-13:30:00.234|
41=ORD12345|11=CXL12346|55=AAPL|54=1|60=20260505-13:30:00.234|10=201
# 41=OrigClOrdID (order being cancelled), 11=ClOrdID (new id for the cancel itself)

# FIX 5.0 SP2 — OrderCancelReplaceRequest / modify (35=G)
8=FIXT.1.1|9=210|35=G|49=PARTICIPANT|56=VENUE|34=44|52=20260505-13:30:00.345|
41=ORD12345|11=RPL12347|55=AAPL|54=1|38=50|40=2|44=185.40|60=...|10=199
# 41=OrigClOrdID, 11=new ClOrdID, 38/44 = replacement qty/px

# OUCH 4.2 — Enter Order (binary, 49 bytes)
struct {
  char  msg_type;        // 'O'
  char  order_token[14];
  char  side;            // 'B' / 'S' / 'T' / 'E'
  uint32 shares;
  char  stock[8];
  uint32 price;          // pennies × 10000
  uint32 tif;
  ...
};

# OUCH 4.2 — Cancel Order ('X')
struct {
  char  msg_type;        // 'X'
  char  order_token[14]; // existing order to cancel
  uint32 shares;         // shares-to-cancel (0 = full cancel)
};

# OUCH 4.2 — Replace Order ('U')  — replace-down / modify-in-place
struct {
  char  msg_type;        // 'U'
  char  existing_token[14];   // order being replaced
  char  replacement_token[14];// new token; replace can only reduce shares / amend px-down
  uint32 shares;
  uint32 price;
  uint32 tif;
};

Cancel & modify are first-class verbs. OUCH Replace is replace-down/modify-in-place semantics (you can shrink qty or amend price, not grow size and keep priority); FIX 35=G carries the same intent. Cancel-on-disconnect (drop the FIX/OUCH session → venue cancels that session's resting day orders) and mass-cancel (per-session / per-account panic cancel) are part of the order-entry contract, not an add-on — the cancel_replace_ratio anomaly signal in the Knight shadow-risk path watches exactly this verb traffic.

**Why FIX and OUCH:** OUCH is binary, fixed-length, ~50 B/msg — built for HFT round-trip in single-digit µs. FIX is text, variable-length, but every OMS speaks it; we accept it for slower flow at relaxed budgets (~50 µs gateway-decode vs ~3 µs for OUCH).

ITCH 5.0 incremental MD — every add/modify/exec/cancel publishes a fixed-length record (typically 30–60 B) on UDP multicast group A and B. Sequence numbers are per-shard.

06Data model

What gets stored, and what is it looked up by?

Sequencer journal record (canonical event, ~80 B):

fieldtypenotes
seq_numu64monotonic, per-shard
ingress_tsu64sequencer-stamped, ns since UTC epoch
event_kindu8NewOrder / Cancel / Modify / Halt / Auction / etc.
participant_idu32account-level
symbol_idu32mapped from stock symbol via static dict
sideu8buy/sell
qtyu32
priceu64price × 1e8 (no float)
order_refu64client-assigned
crcu32CRC32C, validated on every replay

Matcher book (hot, in-RAM, per-shard):

book[symbol_id] = {
  bids:  TreeMap<price, IntrusiveDList<Order>>   // price-level FIFO, descending
  asks:  TreeMap<price, IntrusiveDList<Order>>   // ascending
  by_ref: HashMap<order_ref, Order*>             // O(1) cancel/modify
  state:  AUCTION | CONTINUOUS | HALTED | LULD_PAUSE
}

No tombstones, no soft-delete. Cancels remove the node from the intrusive list. The journal IS history; the book is a projection.

Surveillance tickdb (kdb+):

trade: ([] sym:`g; time:`p; price:`f; size:`j; buyer:`g; seller:`g)
quote: ([] sym:`g; time:`p; bid:`f; ask:`f; bsize:`j; asize:`j)
order: ([] sym:`g; time:`p; ref:`j; side:`c; price:`f; size:`j; status:`c)

Partitioned by date, parted by sym; LZ4 on disk; HDB queries hit at 1–10 GB/s/node.

Database choices — recommended:

  • Sequencer + journal: Aeron Cluster (Raft over reliable UDP, off-heap, deterministic, single-digit µs commit). Chronicle Queue Enterprise is the commercial alternative — broker-less mmapped, used at every tier-1 bank. Don't reach for Kafka — Kafka's commit latency p99 is 1–10 ms, three orders of magnitude over our budget.
  • Drop-copy: Kafka (KRaft, RF=3, min.insync.replicas=2). 100ms p99 is fine here; partitioning by clearing-account gives stable consumer groups for every clearing firm.
  • Tick store: kdb+/q — there is no real alternative at exchange scale. ClickHouse is competent for surveillance queries but the ingest path won't keep up at peak on a single tickerplant. ClickHouse is fine as a secondary surveillance layer.
  • Cold archive: S3 + Glacier Deep Archive with Object-Lock Compliance mode. Counts as SEC 17a-4(f) WORM media (the SEC issued a 2022 interpretive release blessing Object-Lock).

Why not Postgres? Sub-µs commit is not on the table. Postgres replication is async by default (loses RPO=0); even with sync replication, p99 commit is 5–50 ms. Wrong tool.

07High-level design

Which components handle a request, and in what order?

Architecture summary:

  1. Participant → Colo Fabric: FIX/OUCH session over TCP. Cross-connect into the venue's colocation cage. The colo fabric (Arista 7130 / Cisco Nexus 3550-F) hardware-timestamps at line rate and steers per-session traffic onto a VLAN bound to a specific gateway pool.
  2. FIX/OUCH Gateway: Java + Aeron + Solarflare/OpenOnload (ef_vi) on isolated cores. Decodes SBE/FIX/OUCH, applies session-level throttle/credit, hands the canonical event to in-process risk.
  3. Pre-Trade Risk: Inline (in-process) — fat-finger, max position, max notional, self-match prevent, kill switch, restricted list. Implemented as branchless checks against an mmapped credit table. Reg SCI / Rule 15c3-5 require this in path; Knight Capital is the cautionary tale.
  4. Sequencer (Aeron Cluster, Raft): The source of truth. 3 nodes per shard, sync replication, NVMe journals. Every accepted event is assigned a monotonic seq# + sequencer-stamped timestamp. RPO=0, majority quorum commit.
  5. Matcher (single-writer per shard): LMAX Disruptor consumes ordered events from the sequencer. Pure function (snapshot, ordered events) → snapshot'. One thread per symbol-shard, pinned to an isolated core. Hot-standby tails the same Raft log and is bit-identical at every committed offset.
  6. Market Data Publisher: Reads matcher output, emits ITCH/EOBI on UDP multicast A and B feeds. No conflation on incremental.
  7. MD Snapshot / Recovery: TCP unicast service for participants who detect a gap. Replays from the sequencer journal at a requested seq#.
  8. Drop-Copy: Kafka stream of every fill/cancel/modify. Subscribers: clearing, OMS, regulator. p99 enqueue→consume < 100 ms.
  9. Tickdb (kdb+): Tickerplant tails the sequencer journal; RDB holds today; HDB partitions by date+sym. Fuels surveillance and CAT.
  10. Post-Trade Processor: Reads journal end-of-day. Generates DTCC/NSCC settlement files (T+1) and CAT reports (8:00 AM T+1). These external hand-offs (CAT submission e14, settlement-file delivery e15, reconcile poll e32) are not 2PC — you can't run a shared prepare/commit with FINRA CAT or a CCP across an org boundary. The canonical pattern is idempotent at-least-once submission keyed by (date, shard, seq-range) + a T+1 ack-poll reconcile (re-runnable batches, dedupe on the regulator/CCP side).
  11. Cold Archive (S3 + Glacier): Daily journal segments → 30 days S3 Standard → Glacier Deep Archive with Object-Lock Compliance for 7 years.
  12. Ops Control / Kill Switch: Halts, throttles, kill switches enqueued through the sequencer (no out-of-band writes to the matcher). Two-person rule for venue-wide halts.
  13. PTP Time: Dual GPS-disciplined grandmasters → boundary clocks → IEEE 1588v2 PTP to every host with hardware NIC timestamping. Auto-halt rule fires at 50 µs skew.
  14. External — Clearing & CAT: DTCC/NSCC/OCC for clearing, FINRA CAT for the consolidated audit trail.

Hot path on order entry: Participant → Colo Fabric → FIX/OUCH Gateway → Risk → Sequencer (Raft commit) → Matcher → MD Publisher → multicast back to all participants. Drop-copy fans out async. Tickdb tails async. Cold archive ships EOD.

Hot path latency budget (wire-to-wire):

Stagep50p99
NIC ingress + decode2 µs4 µs
Pre-trade risk3 µs5 µs
Sequencer Raft commit (parallel)4 µs8 µs
Match (single-writer)5 µs10 µs
MD encode + multicast TX6 µs10 µs
Misc (queue handoff, NIC TX)5 µs13 µs
Total (ingress → MD egress)25 µs50 µs

08Deep dives

Which part breaks first, and what do you do about it?

1. Determinism — the load-bearing invariant. The matcher is a pure function. Replaying the journal must produce the same book. That means: no System.currentTimeMillis() inside the match (timestamps come from the sequencer record); no HashMap iteration in hot paths (use sorted intrusive lists); no allocator-dependent behaviour (object pools, off-heap); no JIT warmup leaks (warm in production-equivalent test prior to session open, then never recompile during RTH); single-threaded match loop. Surveillance, MD, drop-copy are projections of the log — they catch up async; they're never canonical.

2. Knight prevention. Risk is in-process, on the gateway core, in front of the sequencer. Every order touches it. Bypass is impossible by topology — there is no edge from gateway → sequencer that doesn't traverse risk. Deploys are bit-identical (verified by binary hash), feature flags are retired aggressively (max 30 days from introduction), and a parallel "shadow risk" runs on drop-copy detecting cancel/replace ratio anomalies in real time. The kill switch is wired to the sequencer (halts enqueued in-band) AND to a hardware path on the colo fabric (cuts session-level traffic at line rate).

3. Hot-standby = bit-identical (the FB IPO 2012 lesson). The secondary matcher is the same binary, same JVM, same kernel, same NIC firmware as the primary. It tails the same Raft log. At every committed offset its book hash equals the primary's book hash — verified by a continuous parallel "shadow matcher" that emits a book_state_hash_mismatch alert if they ever diverge. Failover is "stop primary, promote secondary, resume" — never "primary down, run different code path." Tokyo Arrowhead 2012 (memory bit-flip + failover to a divergent secondary) and Nasdaq FB IPO 2012 (secondary skipped a cancel-validation path) are the canonical "state alone is not enough" disasters.

4. Multicast A/B + recovery. Every market-data event is published twice on disjoint network paths. Subscribers run a sequence-number-aligned merge: A or B, whichever arrives first; gap-detect on either feed; pull missing range from the unicast snapshot/recovery service. Nasdaq's MoldUDP64 and CME MDP 3.0 are the reference designs; we follow their playbook.

5. PTP + halt rule. Dual GPS antennas on diverse roof paths, plus an independent third source (eLoran or peer-venue PTP via dark fiber) so a roof-level antenna disturbance can't take both grandmasters at once. Two LANTIME M1000 grandmasters, OCXO holdover spec ≥ 24 h within 1 µs. Boundary-clock switches (Arista 7130) distribute PTP IEEE 1588v2 in authenticated mode (Annex K / IEEE 1588-2019 prong) — auth: mtls on the PTP edges represents the network-layer key authentication, not TLS itself. Internal SLO is ±20 µs vs UTC; auto-halt rule fires at 50 µs (half MiFID II RTS 25 budget). Reportable to ESMA if the venue books any execution above 100 µs skew.

Pre-market PTP-clean gate. The auto-halt rule only matters if state == CONTINUOUS. Pre-market (04:00–09:30 ET) the matcher is idle/AUCTION-warming, so PTP loss does not trigger the rule — but the tickerplant ingests with bad timestamps for hours. We add a pre-market PTP-clean gate: the auction open at 09:30:00 is blocked if ptp_offset_ns:p99 > 25_000 over the prior 60 minutes. The gate is wired through ops-control → sequencer (in-band halt event), so a clock-degraded venue cannot open. PTP-recovery re-stamp policy: post-recovery timestamps are re-stamped with (sequencer.ingress_ts, ptp_recovered_marker) so surveillance can distinguish "during outage" from "post-recovery" rows.

6. Auction → continuous transition (the FB IPO race). At 09:30:00.000 the matcher (a) quiesces inbound (sequencer queues but does not deliver new orders), (b) freezes the auction book, (c) runs the cross deterministically against the frozen book, (d) emits prints in sequencer order, (e) flips state to CONTINUOUS, (f) un-quiesces. The race in 2012 was "detect new orders during validation re-run" — the fix is the freeze step plus single-writer.

7. Microbursts. A 200K pps burst over 50 ms blows past kernel TCP/IP stacks. Solarflare/OpenOnload (or AMD-Xilinx X3522 / DPDK) puts userspace in the receive path with NIC ring buffers in hugepages. Per-session token-bucket throttle at the gateway absorbs participant-side bursts; per-shard sequencer queue depth is the load-shedding signal.

8. Halt blast-radius. Symbol halt at 14:32:17 → ops-control signs the halt event → enqueues via sequencer → matcher consumes the halt event in order on that shard → match loop pauses → in-flight orders are cancelled or held per LULD rules → MD publishes halt → gateway rejects new orders for that symbol with a session-level reject. Order-of-operations is the halt — every step is an event in the log, replayable, audit-evident.

09Trade-offs

What did this design cost, and what breaks at 10×?

What we explicitly chose, with the alternative we rejected:

  • Aeron Cluster (Raft) over Kafka. Kafka p99 commit is 1–10 ms. We need ≤8 µs. Kafka loses by three orders of magnitude.
  • Single-writer per shard over multi-threaded. Multi-threaded match is faster per core but non-deterministic on replay. We trade throughput we don't need (5M msgs/s/shard ≫ peak) for determinism we cannot do without.
  • In-process risk over sidecar. Sidecar adds a network hop and a serialization. At µs budgets, in-process is the only option. Reg SCI requires risk in path; physically and administratively co-locating it removes the bypass risk.
  • Single-region active-passive over multi-region. Latency arbitrage forbids cross-region matching; HFT firms would race the speed of light between regions. Same-metro hot-standby with synchronous Raft replication is the industry consensus.
  • UDP multicast over TCP for MD. Multicast scales to thousands of subscribers at one publish cost. TCP would require N TCP sessions per event. Trade-off: lossy → A/B + snapshot recovery covers it.
  • Event-sourced over outbox. The journal IS the source of truth; the matcher is a projection. Outbox would imply a separate system-of-record (e.g. a SQL DB) and a stream — we don't have that, the log is the DB.
  • kdb+/q over ClickHouse. ClickHouse is excellent but kdb+ is the standard surveillance / TCA stack; every quant team and every regulator interface is built around it. Network effects beat the OSS license.

Primary sources

  • LMAX Disruptor paper
  • Martin Fowler — The LMAX Architecture
  • Aeron Cluster Raft consensus
  • Nasdaq TotalView-ITCH 5.0
  • Nasdaq OUCH 4.2
  • SEC Release 34-70694 (Knight Capital)
  • SEC press release 2013-95 (Nasdaq Facebook IPO)
  • MiFID II RTS 25 (clock synchronisation)
  • SEC Rule 613 / CAT NMS Plan
  • CFTC-SEC Joint Report on the May 6, 2010 Flash Crash

Now defend it

Reading a design is not the same as being able to hold one under questioning. The workspace asks the same questions an interviewer would, and the simulator disagrees with you when the diagram does not support the claim.

Work Stock Exchange / Order Matching yourself

More in Transactions, Concurrency & Money

Correctness when two writers collide and money is involved: serializability, two-phase commit versus sagas, hold-then-confirm, single-writer matching, and the databases that give you external consistency.