AI Agent Platform
Worked solution

AI Agent Platform — a worked solution

Long-running multi-step LLM agents — durable workflow + sandboxed execution + LLM gateway. Brain / hands / state are independently replaceable. The agent run is a workflow, not a request.

Try it yourself first.

You will remember almost none of this if you read it cold. The workspace walks the same 10 stages and runs the architecture you draw through a simulator, so you find out where your version breaks before you see ours.

Open the AI Agent Platform workspace

The problem

Build the platform that powers long-running multi-step LLM agents — Devin, Manus, OpenAI Codex agents, Cursor agents, Claude Code background agents, Replit Agent, Lindy. Users describe a task; the platform spawns an agent that plans, calls an LLM, calls tools, executes code in a sandbox, captures results, and iterates — sometimes for minutes, sometimes for hours, sometimes (Cursor's Solid→React migration) for three weeks. The architectural pressure here is not LLM inference (that's an external dependency) — it's durability, isolation, and cost containment under failure.

Three load-bearing concepts:

  1. An agent run is a durable workflow, not a request. Brain (LLM gateway) / hands (sandbox) / state (orchestrator + history) are independently replaceable. When the sandbox host dies 90 minutes into a 2-hour Devin task, the workflow survives because Temporal's history is the source of truth. (Cognition's published architecture; Temporal's Replit case study.)
  2. KV-cache-aware routing is not optional. Manus's blog calls KV-cache hit rate "the single most important metric for a production-stage AI agent" — Anthropic's prompt-caching read price is 10% of base, so a 90% cache hit rate cuts cost by ~80%. A naïve OpenAI-protocol-compatible passthrough cannot do this; the gateway must understand prefix locality.
  3. Cost runaway is enforcement, not alerting. The published $47K LangChain incident (four agents in an unintended infinite loop, 11 days, $47K bill) is the canonical retro. Per-tenant kill-switches must fire at minute granularity, not at the daily billing job. Replit users have reported $30/hr → $360/day before their 2025 controls landed.

The reference architecture

Reference architecture for AI Agent Platform: 17 components — Client, CDN · SSE Edge, API Gateway, Auth / Identity, Service · Run Orchestrator (Write Plane), Service · Stream Reader (Read Plane), Service · LLM Gateway, Service · MCP Tool Router, Worker · Sandbox Pool, Coordinator · Temporal Cluster, NoSQL DB · Workflow History, SQL DB · Run Metadata + Idempotency + Vector Memory, Stream · Event Bus, Worker · Cost Aggregator + Kill-Switch, Object Store · Transcripts + Snapshots, External · LLM Providers, External · MCP Servers — connected by 27 flows.ClientclientCDN · SSE EdgeCloudflare Workers + WAFAPI GatewayEnvoy + Cloudflare WAF …Auth / IdentityOAuth 2.1 PKCE + JWKS r…Service · Run Orchest…Go + connect-go to Temp…Service · Stream Read…Go + Kafka consumer + S…Service · LLM GatewayRust + Tokio; LiteLLM-p…Service · MCP Tool Ro…TypeScript + MCP SDK; O…Worker · Sandbox PoolFirecracker microVMs + …Coordinator · Tempora…Temporal Server 1.27 (f…NoSQL DB · Workflow H…Cassandra 5.0 (Temporal…SQL DB · Run Metadata…Postgres 16 + Citus + p…Stream · Event BusKafka 3.7 (KRaft mode),…Worker · Cost Aggrega…Go + Kafka consumer + P…Object Store · Transc…S3 + Object Lock + Glac…External · LLM Provid…Anthropic Claude API, O…External · MCP ServersAnthropic-spec MCP (OAu…
17 components, 27 flows. A dashed line is an asynchronous hop. This is the reference design, not the only one that works.

What each component is for

Client

Three concrete callers share this entry point: (a) an IDE / desktop agent (Cursor, Claude Code, Devin web UI) sending POST /v1/runs, (b) a CLI tool subscribing to GET /v1/runs/:id/events for a long-lived SSE token stream, and (c) a programmatic client (CI, batch job) polling for terminal status. Receives webhook callbacks at a registered URL when a run finishes asynchronously.

Why it exists. Considered modeling each surface as a separate node; rejected because all three share the same auth, idempotency, and SSE-resume contracts. Splitting only matters at the SDK layer (different language SDKs, retry helpers); the system-level failure modes are identical.

When it fails. Naïve clients reconnect SSE without Last-Event-ID and re-render duplicated tokens — looks like a model regression to users. Detection: distinct event_ids/run p95 > 1.05 alerts SDK eng. OpenAI Assistants API has the documented 'stream interruption' bug in this shape (community.openai.com).

CDN · SSE EdgeCloudflare Workers + WAF

Terminates TLS, runs the WAF ruleset (OWASP Top-10 + agent-platform rules: rejects Authorization-header replay attempts, fingerprints scraper UAs), absorbs SSE token streams (long-lived TCP) so the origin tier doesn't sit on 50K+ open sockets, enforces a per-IP global token-bucket BEFORE the gateway. POST requests pass through to the gateway; GET /v1/runs/:id/events terminates here and proxies upstream via Cloudflare's persistent connection pool.

Why it exists. Considered terminating SSE at the gateway; rejected because at peak (~50K concurrent runs × 1 SSE stream each = 50K open connections) Envoy's connection pool exhausts and one bad deploy starves the entire fleet. Edge SSE termination buys us connection multiplexing and, more importantly, a WAF / abuse layer that doesn't share fate with the agent control plane. Cloudflare Workers + the global anycast reach is what makes this affordable; building it in-house is months of engineering for no product gain.

When it fails. Edge degrades → SSE waves disconnect; runs continue server-side (the durable-workflow property holds) but UI shows frozen. Detection: cf_status_5xx_rate, sse_disconnect_rate, gateway_active_streams. Mitigation: client SDK falls back to the long-poll mode of GET /v1/runs/:id/events?mode=poll&since=event_id (event-delta semantics) until SSE recovers. The OpenAI Assistants API stream-interrupt class of bug lands here when the CDN tier is the culprit.

API GatewayEnvoy + Cloudflare WAF + per-tenant token bucket

L7 ingress for every public agent-platform call: POST /v1/runs (start), POST /v1/runs/:id/messages (mid-run user message), POST /v1/runs/:id/cancel, GET /v1/runs/:id, GET /v1/runs/:id/events (SSE — passes through to stream-svc). Validates the JWT against auth, enforces per-tenant token-bucket rate limits BEFORE forwarding so a runaway client can't poison the orchestrator, copies the Idempotency-Key header verbatim onto run-svc, and stamps a request-correlation header (W3C Trace Context) so a single agent run is traceable end-to-end through Temporal activities and tool calls.

Why it exists. Considered baking auth + rate limit inline in run-svc; rejected because (1) a 5xx in run-svc would skip the rate limiter and let the storm reach Temporal — and Temporal's frontend is a finite resource (4 096 history shards is the published ceiling per cluster), (2) per-tenant rate limits are tenancy data we'd otherwise ship to every run-svc replica, and (3) gateway 401/403 responses are constant-time, which run-svc isn't. Centralizing authn behind the gateway means run-svc trusts the signed envelope without a round trip. Stripe's 2019-07-10 retro is the 'rate limit in the wrong layer' anti-lesson.

When it fails. Gateway pool exhaustion masks as 'product is slow' but is upstream load. Detection: gateway_active_streams (50K SSE pass-through is the line), downstream_5xx_rate, per-tenant 429 rate. Mitigation: drain mode + per-zone failover; circuit-break a misbehaving tenant rather than 503'ing everyone. The 'one tenant in a tool loop' pattern (the $47K LangChain incident) presents here first — a tenant's run rate spikes past their bucket, the bucket holds, the cost-worker downstream eventually flips the kill-switch.

Auth / IdentityOAuth 2.1 PKCE + JWKS rotation + Postgres key store

Validates the Authorization header on every public call: hashed key lookup against a sharded Postgres of API keys for service-account auth, OAuth 2.1 + device binding for human users (CLI / IDE installers), JWKS rotation for partner integrations. Returns a signed (merchant_id, tenant_id, scopes, daily_budget_remaining_usd, concurrent_run_quota) envelope back to the gateway, valid 60 s. The same envelope is consumed by run-svc, llm-gw, and mcp-router so each can make scope decisions without an extra hop.

Why it exists. Considered baking authn into run-svc; rejected because (1) a CVE in the run-svc image should not break authn for everyone, (2) restricted scopes are a separate product surface (read-only viewer keys, billing-only keys), (3) the lookup hits a 100M-row keys table that doesn't belong on the agent hot path. Centralizing authn behind the gateway with a 60 s scope cache means the gateway rate-limits before any expensive work, and run-svc trusts the envelope.

When it fails. Auth down → all writes 503 (fail-closed); reads with cached scopes survive 60 s then 401. Detection: auth_5xx_rate, scope_cache_miss_rate. Cloudflare R2 March 21 2025 (credential rotation pushed to dev not prod, then old creds deleted; 100% writes / ~35% reads failed for >1 h) is the cautionary tale for cert/JWKS rotation playbooks.

Service · Run Orchestrator (Write Plane)Go + connect-go to Temporal; Brandur-style idempotency state machine

Hosts every mutating run-lifecycle endpoint: POST /v1/runs (create), POST /v1/runs/:id/messages (signal a mid-run user message), POST /v1/runs/:id/cancel (signal cancel), POST /v1/runs/:id/approve (signal HITL approval). For POST /v1/runs: hashes (tenant_id, request body) into a request_fingerprint, runs the Brandur-style idempotency state machine (claim → started → workflow_signaled → finished), commits a SERIALIZABLE Postgres transaction that writes the run_meta row + the idempotency row in one shot, then issues a Temporal StartWorkflowExecution gRPC call. Returns 201 with run_id and the SSE URL in <500 ms; the agent loop runs asynchronously inside Temporal.

Why it exists. Considered a single combined read+write service. Splitting Run Orchestrator (write) from Stream Reader (read) lets each scale on its own load shape (write QPS ≈ 0.05× read SSE concurrency), isolates blast radius (an SSE storm can't crater run starts), and makes per-endpoint rate limits cleaner. The idempotency state machine MUST live here rather than as a sidecar — Brandur's writeup proves correctness only when state transitions interleave with the business write in the same DB transaction. Considered putting StartWorkflowExecution inside the same SERIALIZABLE txn (2PC across Postgres + Temporal); rejected because Temporal rejects writes during partial-failure and a 2PC commit hold-time would defeat the <500 ms SLO. The outbox pattern (commit run-meta + outbox row, then Temporal-relay drains) is the load-bearing alternative — cleaner failure semantics, +20 ms latency.

When it fails. Idempotency store down → fail-CLOSED 503 (cannot accept a run start without a claim row). Temporal frontend down → run-meta commits, outbox-relay retries SignalWorkflowExecution with exponential backoff; client sees 202 Accepted with a status URL once run_meta exists. Detection: claim_p99 < 50 ms, idem_lock_wait_ms p99 < 100 ms, outbox_relay_lag_seconds < 10. The fail-closed posture is the same Stripe took in 2019-07-10: better to 503 than to silently lose a run start.

Service · Stream Reader (Read Plane)Go + Kafka consumer + S3 transcript reader

Hosts GET /v1/runs/:id/events (SSE) and GET /v1/runs/:id (terminal status snapshot). For SSE: subscribes to the run's Kafka partition (key = run_id) and pushes token / tool-call / status events to the client as Server-Sent Events with monotonic event_id. Honors Last-Event-ID header on reconnect — replays from the partition offset matching that event_id, falling back to S3 transcript archive when the requested offset is older than Kafka retention (24 h hot tier). For status snapshot: cache-aside through Redis (tenant-scoped); falls through to run-meta and (if the run is finished) S3 transcript.

Why it exists. Considered serving SSE from run-svc; rejected because SSE concurrency is ~1 000× run-start concurrency (each run streams for its full duration). A stuck SSE connection pool would back-pressure new runs. Considered making the cache the source of truth for transcripts; rejected because the audit trail must persist for 90 days regardless of cache state. Stream-svc holds NO state — it's pure cache + Kafka consumer + S3 reader. Resumable streams (Last-Event-ID) are non-negotiable: the OpenAI Assistants API's stream-interrupt bug pattern is what happens when the read plane couples run lifecycle to socket lifecycle.

When it fails. Cache cold → status snapshot reads hit run-meta replica → replica saturates → switch to degraded-mode that returns last-known status with a 'may be stale' header. SSE backbone (Kafka) down → existing SSE streams die; clients reconnect with Last-Event-ID; S3 fallback reads keep working but at higher latency. Detection: sse_disconnect_rate, kafka_consumer_lag_ms, sse_active_streams, last_event_id_mismatch_rate.

Service · LLM GatewayRust + Tokio; LiteLLM-protocol-compatible upstream; KV-cache-aware sticky router

Single egress point for every LLM call. Three load-bearing responsibilities: (1) Routing — choose provider (Anthropic / OpenAI / Bedrock / Vertex) per call based on (tenant_id, model_id, region, current error-rate per provider, KV-cache locality); failover when a provider's observed error rate exceeds 30% (NOT when its status page is red — Anthropic Aug 2025 misroute lesson). (2) Caching — sticky-pin requests with the same prompt prefix to the same upstream pod so vLLM/RadixAttention prefix caching hits; for Anthropic, set the anthropic-beta: extended-cache-ttl-2025-04-11 header and cache_control: { ttl: "1h" } because Anthropic silently dropped the default to 5 min in April 2026. (3) Quota enforcement — pre-call hard cap per (tenant, day) and per (run, total): rejects with 402 Payment Required if the kill-switch flag is set on the tenant by cost-worker; rejects with 429 if the per-run token cap is exceeded. The kill-switch is the load-bearing piece: an *alert* is not enforcement.

Why it exists. Considered a thin OpenAI-protocol-compatible passthrough (LiteLLM, Portkey). Rejected because routing has to be cache-aware — Manus's published lesson is that KV-cache hit rate is the single most important production metric, and an OpenAI-protocol passthrough cannot do prefix-locality routing because it has no awareness of the upstream's cache topology. Considered making each service call its own provider directly. Rejected because (1) cost arithmetic would be impossible without a single egress meter, (2) cross-service circuit-break decisions need a shared state, (3) a single PCI-style 'all LLM tokens flow through here' boundary is what makes per-tenant quota auditable.

When it fails. All providers degraded → circuit-break per (provider, region, model); fall back to next-best provider if cross-provider tool-call schema can be normalized (kept normalized in the orchestrator, not provider-specific). Cache-aware routing tier saturates → degrade to round-robin; cache hit rate drops 85% → 60%; cost rises ~3×; pages oncall on cost_per_run_dollars SLO breach (NOT on provider_5xx, which would miss the gray failure). Anthropic Aug 2025 misroute (Sonnet 4 → 1M-context servers) is the canonical 'gateway routes by capacity tier, not by SLI' anti-pattern.

Service · MCP Tool RouterTypeScript + MCP SDK; OAuth 2.1 PKCE; per-tool circuit breaker

Single egress point for every tool call (MCP servers + first-party tools). Validates the (run_id → allowed_tool_set → required_oauth_scope) capability triple before any call lands at a third-party server, runs a per-tool circuit breaker that opens after 50% errors over 30 s, terminates the OAuth 2.1 flow with PKCE for user-bound tools, caches tools/list per server (with cache_tools_list=true because regenerating the tool catalog every call is the published MCP scaling foot-gun). Returns the tool result to the caller, redacted by the per-tenant output classifier (SecretRedactor + PIIRedactor) because tool output is the prompt-injection attack surface.

Why it exists. Considered letting the orchestrator activity workers call MCP servers directly. Rejected because (1) the prompt-injection 'lethal trifecta' (private data + untrusted content + external comms — Simon Willison) is broken exactly here: a single egress point can scrub tool output before it reaches the next LLM call, (2) per-tool kill-switches must work fleet-wide, (3) the EscapeRoute Anthropic filesystem MCP CVE-2025-53109/53110 (path-traversal) is the kind of bug that needs a defense-in-depth boundary the agent runtime CANNOT bypass. Considered making capability scoping a sandbox-pool concern (egress proxy in the microVM); rejected because some tool calls happen from orchestrator activities outside any sandbox.

When it fails. Per-tool circuit open → that tool returns 'unavailable' to the agent; the model sees a clean error and pivots in the prompt. Classifier stalls → fail-CLOSED on tool output (better to drop a tool call than to leak). Detection: per-tool circuit_open_rate, classifier_p99, oauth_token_refresh_failure_rate. The Anthropic April 23 2026 Claude.ai MCP-applications outage (00:41-02:09 UTC, v2.1.120 force-rolled-back) is the precedent — a bad MCP layer rollout starves agents fleet-wide, but if circuit breakers fire per-tool, only impacted tool flows degrade.

Worker · Sandbox PoolFirecracker microVMs + KVM + virtio-net + memory-snapshot fork

Hosts the per-task Firecracker microVMs that run agent-generated code, headless browsers (Chromium + Playwright), and untrusted tool-output processing. Each run gets a dedicated microVM forked from a warm snapshot (~150 ms with Firecracker per E2B's published numbers; ~10× faster than fresh boot). Inside the VM: 2 vCPU / 1-2 GB RAM, ext4 root, a unix-domain-socket egress proxy that enforces the per-tenant capability token, /workspace mounted from a per-run scratch volume. Acts as a Temporal activity worker for bash, code_exec, browser_action, read_file, write_file activities. On task end (or eviction), the microVM is snapshotted to S3 and torn down.

Why it exists. Containers (runc) share a kernel, and Cognition's published 'Building Cloud Agents' post documents the exact failure: 'a single compromised session can reach every other container's filesystem.' Firecracker microVMs (KVM hardware virtualization) are the production isolation primitive every coding-agent vendor (E2B, Cognition, Vercel, Devin) converged on. Considered gVisor (user-space kernel, weaker but cheaper cold-start) — chose Firecracker because runc Leaky Vessels (CVE-2024-21626) and the Nov 2025 trio (CVE-2025-31133/52565/52881) prove that container-class isolation is too thin for arbitrary agent-generated code. Considered always-on warm pool only — the published cost crossover with fork-from-snapshot is around 10 sandboxes/sec arrival rate (E2B blog); we run hybrid (warm pool absorbs steady-state, fork-from-snapshot absorbs bursts).

When it fails. Warm pool exhaustion → new runs queue; sandbox-acquire p99 climbs from sub-second to tens of seconds. Detection: sandbox_queue_depth, warm_pool_size, fork_p99. Mitigation: spillover to cold-start tier + emergency host autoscale (autoscale lag 2-3 min). The E2B Mar 9 2026 'Sandbox management degraded' incident is the precedent. Sandbox escape via runc Leaky Vessels-class CVE → quarantine the affected host, reroute to the rest of the pool (1/256 blast radius), forensics on the host. Anthropic's filesystem MCP EscapeRoute (CVE-2025-53109/53110) is the closest agent-platform-specific precedent.

Coordinator · Temporal ClusterTemporal Server 1.27 (frontend + history + matching + worker, multi-cluster replication)

Owns every agent run as a durable workflow keyed by run_id. The workflow code is deterministic Go that drives the agent loop: plan → call LLM (activity) → call tool (activity) → update memory (activity) → repeat. LLM and tool calls are non-deterministic Activities with at-least-once semantics, idempotency keys derived from activity_id, and exponential-backoff retries (5 attempts default, configurable per activity). Workflow history is event-sourced into Cassandra (the workflow-history node) and replayed on any worker restart. Signals (cancel, user-message, HITL approve) are first-class — they're how the read plane and the user influence the run without breaking determinism. Pointer State pattern: large activity payloads (full LLM responses, tool results) are written to S3 and only an S3 pointer lands in the history (99.8% payload reduction per AzGuards' published guidance — keeps history shards small).

Why it exists. The single insight that unlocks this whole design: an agent run is a durable workflow, not a request. LangGraph 1.0 ships checkpointing per super-step, but Diagrid's published critique 'Checkpoints Are Not Durable Execution' shows it lacks at-least-once activity semantics, exponential backoff, and signal-based interrupts that Temporal provides. Cognition's Devin runs on Temporal; Replit Agent 3 runs on Temporal (case study published); OpenAI Codex web agent runs on Temporal. Considered DBOS (Postgres-resident workflow state) and Restate (HTTP sidecar); both are viable, but Temporal's Cassandra-backed history shards and proven 4 096-shard cluster ceiling are what we need at 50 K concurrent runs. The non-trivial trade-off accepted: workflow code MUST be deterministic — every replay produces the same control flow given the same history. This forces all I/O into activities, all randomness into Workflow.NewRandom, all time into Workflow.Now — and any new agent prompt version must gate via Workflow.GetVersion or it explodes on replay (the non_deterministic_error pager).

When it fails. History shard quorum loss → workflows on those shards stall (read+write block); other shards continue. Frontend down → run-svc StartWorkflowExecution returns 503; outbox-relay retries; user sees brief delay. Worker pool down → activities don't dispatch; in-flight workflows block on activity ScheduleToClose. The most insidious failure: non_deterministic_error after a deploy where workflow code changed without a Workflow.GetVersion gate; in-flight workflows fail-closed on replay. Detection: temporal_workflow_task_failed{cause='non_deterministic_error'} > 0 is a P1 page; shard_lock_p99, history_event_persist_p99.

NoSQL DB · Workflow HistoryCassandra 5.0 (Temporal default backing store)

Cassandra cluster that backs Temporal's workflow history. Two main tables: executions (per-run pointer to current history slice) and history_node (the append-only event log per workflow_id). Every workflow event — workflow started, activity scheduled, activity completed, signal received, timer fired — is appended here. The event log IS the source of truth: any worker can replay a workflow from event 0 and arrive at the same state. Partition key = workflow_id (run_id); clustering key = event_id (monotonic per workflow).

Why it exists. Considered Postgres for the history (Temporal supports it for small clusters); rejected because at 50 K concurrent runs × ~500 events / run / hour = 25 M events/hour, Postgres' WAL fan-out becomes a bottleneck around 5-10 K writes/sec. Cassandra's leaderless quorum scales horizontally on writes and partitions cleanly by workflow_id. Considered FoundationDB (some Temporal forks use it); rejected because Cassandra is the documented production-default with the largest community of operators. Considered S3 for history (event-sourced into object store); rejected because per-event read latency is fatal at ~500 events/run replay.

When it fails. Quorum loss (2 of 3 AZs partitioned for one shard's replicas) → workflows on that shard halt; W cannot form, R cannot form. Detection: temporal.persistence.error_rate, cassandra.write_timeout, cassandra.read_timeout. Monzo 2019-07-29 (auto_bootstrap=false on six new nodes returned empty quorum reads) is the canonical retro — the lesson is that scale-up tests must include multi-node bootstraps, not just single-node. Mitigation: hinted handoff queue depth alert at 1 GB; never relax to ONE on writes.

SQL DB · Run Metadata + Idempotency + Vector MemoryPostgres 16 + Citus + pgvector + pgvectorscale

The relational source of truth for: (1) runs table (run_id, tenant_id, workflow_id, status, tags, created_at, terminal_at, billing_cents) — the index every list / search query hits; (2) idempotency_keys table (Brandur-style, 24 h hot + 48 h soft-retain TTL); (3) tenant_meter (running cost-per-day, run-count-per-hour, kill_switch_engaged_at — the row cost-worker writes to); (4) episodic_memory table with pgvector embeddings for cross-run retrieval (tenant-scoped, recall@10 target 0.92). Sharded by tenant_id on Citus (32 shards) so per-tenant queries are single-shard. The SAME engine hosts idempotency and runs because the run_meta INSERT and the idem CLAIM commit in one SERIALIZABLE transaction at run-svc.

Why it exists. Considered DynamoDB for runs+idempotency (single-key conditional writes are perfect for CLAIM); rejected because the runs table has rich query patterns (list by tenant, filter by status, paginate by created_at) that DynamoDB models awkwardly. Considered separating idempotency into Redis; rejected because the idempotency contract requires state transitions atomic with the business write. Considered a dedicated vector DB (Qdrant, Turbopuffer) for episodic memory; for the default scale tier (≤10 M vectors per tenant) pgvector + pgvectorscale gives 471 QPS @ 99% recall on 50 M vectors (TigerData benchmark) and keeps memory co-located with run metadata — the scale-out path to Turbopuffer is documented but not the default.

When it fails. Single shard primary down → 1/32 of tenants see write failures for ~30 s failover. Cross-shard read fan-out for admin / billing queries hits multiple shards in parallel; one slow shard slows the query (head-of-line). Detection: per-shard write_p99, replica_lag_seconds, oldest_unreplicated_lsn, idempotency_lock_wait_p99. Brandur's writeup is explicit: idempotency-store fail-closed is mandatory. Adjacent risk: episodic_memory is a cross-run prompt-injection vector — a malicious tool result that gets embedded and recalled on a future run by the same tenant pollutes downstream context; the output classifier runs on tool returns but a second-stage classifier on memory writes is the load-bearing mitigation (NOT shipped in v1; tracked as accepted residual risk).

Stream · Event BusKafka 3.7 (KRaft mode), 256 partitions, RF=3, acks=all

The async backbone connecting llm-gw, mcp-router, sandbox-pool, run-svc as producers to cost-worker, stream-svc, and the OTel trace sink as consumers. Three logical topics: (1) agent.events.v1 — every coalesced token chunk, llm.call, tool.call, sandbox.lifecycle, run.status; partitioned by run_id (preserves SSE order for stream-svc). (2) cost.meter.v1 — per-call cost samples; partitioned by tenant_id (preserves cost-worker single-writer per tenant). (3) audit.v1 — security-audit events (capability-token-issued, kill-switch-fired, cert-rotation-completed); partitioned by tenant_id, RF=5, immutable. Producers use the idempotent producer with acks=all + min.insync.replicas=2 = exactly-once into Kafka.

Why it exists. Considered direct service-to-service calls (llm-gw → cost-worker via gRPC). Rejected because (1) cost enforcement needs a durable backplane — losing a cost event means the meter under-counts and the kill-switch fails late; (2) SSE fanout to many concurrent stream-svc replicas needs broadcast semantics; (3) the audit topic is the load-bearing piece for compliance — every capability-token issuance and every kill-switch fire is a regulator-readable event. Considered Pulsar (better tenancy isolation); chose Kafka because the operator pool is larger and the published Cassandra+Kafka+Temporal stack is a known-good shape. Considered NATS JetStream; rejected because the durability guarantees aren't acks=all+ISR equivalents at this scale.

When it fails. ISR shrinks below min.insync.replicas → producers block; llm-gw and sandbox-pool back-pressure; SSE token streams pause. Two-broker loss in same AZ stops producers for that AZ. Detection: under_replicated_partitions, isr_shrink_rate, producer_request_latency_p99. Mitigation: 3-AZ broker spread is non-negotiable; per-AZ producer affinity so a single-AZ broker outage is degraded-not-broken.

Worker · Cost Aggregator + Kill-SwitchGo + Kafka consumer + Postgres SELECT FOR UPDATE

Consumes cost.meter.v1 from Kafka, aggregates per-tenant cost in 1-second buckets, writes the running cost into tenant_meter in run-meta. When (cumulative cost / day > tenant_budget) OR (cost / minute > runaway_threshold) OR (no-progress detector — same tool args repeated >5 times in 60 s), engages the kill-switch: writes kill_switch_engaged_at = now(), publishes a control signal to llm-gw via gRPC (cancels in-flight + rejects new calls for that tenant), and pages on-call P1.

Why it exists. Considered enforcing the cap inside llm-gw directly — cumulative state per tenant lives in Redis, every call decrements. Rejected because (1) Redis isn't durable (a flush would lose the meter), (2) cross-replica consistency on the meter is required (32 llm-gw replicas all writing to the meter row would cause hot-row contention), (3) the no-progress detector is a windowed-stream computation that doesn't fit at the gateway. Centralizing into a single-writer-per-tenant cost-worker (Kafka partition by tenant_id ensures this) gives us atomic-in-the-stream meter updates. The $47K LangChain incident retro is the lesson: an open-loop billing alert is not enforcement — the kill-switch must fire at minute granularity, not at the daily billing job.

When it fails. Cost-worker consumer lag > 60 s → meter updates stale → kill-switch fires late → tenant overspends past their cap. Detection: kafka_consumer_lag_seconds (P1 at >30 s), kill_switch_decision_latency_p99, runaway_detected_count. Mitigation: scale-out via cooperative rebalance (no stop-the-world); fallback hard-cap at gateway (1.5× tenant budget) so even a fully-failed cost-worker bounds the loss. Replit users reported $30/hr → $360/day before their 2025 controls landed; this node prevents that.

Object Store · Transcripts + SnapshotsS3 + Object Lock + Glacier tier

Cold-tier and audit-grade store for: (1) per-run full transcripts (every LLM call request+response, every tool call, every sandbox stdout/stderr) keyed runs/{tenant_id}/{run_id}/transcript.jsonl; (2) sandbox memory snapshots from Firecracker for resume-on-crash, keyed sandbox/{run_id}/snap-{seq}.bin; (3) Pointer State payload offloads from Temporal (large LLM responses, tool results) keyed temporal/{workflow_id}/payload-{event_id}.bin; (4) audit.v1 long-term sink with Object Lock (compliance-grade WORM for 7 years). Multi-region replication to a DR bucket; Glacier tier for runs > 90 days old.

Why it exists. Considered keeping transcripts in run-meta (Postgres); rejected because per-run transcript can hit 100 MB on long Cursor / Devin sessions and Postgres TOAST + WAL volume would be catastrophic at scale (AzGuards documented the 99.8% payload reduction from offloading). Considered DynamoDB for transcripts; rejected on cost (S3 is ~25× cheaper for cold reads). The Pointer State pattern (write the heavy payload to S3, only the S3 URL into Temporal history) is what keeps Temporal history shards small and replay fast.

When it fails. S3 regional event → reads fail; resume-on-crash slows (sandbox snapshot unreadable) → microVM cold-starts from base instead of from snapshot. Detection: s3_5xx_rate, snapshot_load_p99, transcript_write_failure. Mitigation: cross-region replication with DR-region active read fallback; Pinecone March 2023 (515 indexes deleted by buggy SQL cleanup) is the cautionary tale for any data-store with bulk-delete capability — Object Lock COMPLIANCE prevents the equivalent on audit data.

External · LLM ProvidersAnthropic Claude API, OpenAI, AWS Bedrock, Google Vertex

Third-party LLM inference. Anthropic Claude (Opus 4.7 / Sonnet 4.6 / Haiku 4.5) for the primary route, OpenAI GPT-5 / o3 for fallback, Bedrock for VPC / data-residency-bound tenants, Vertex for GCP-native customers. Each provider's tool-call format is provider-specific (Anthropic tool_use blocks vs OpenAI tool_calls arrays); the orchestrator stores normalized form so failover is possible mid-run. Streaming responses come back via SSE on the provider side; llm-gw relays into the event-bus.

Why it exists. Considered self-hosting (vLLM + open-weight Llama / Mixtral / Qwen). Rejected for the default canonical because frontier-model quality (Claude Opus 4.7 / GPT-5) is the load-bearing piece for agent reliability, and no open weight matches it as of 2026. The KV-cache-aware llm-d / Nvidia Dynamo path is documented for the vendor-lock-mitigation tier (large enterprises) but not the default. Considered single-provider; rejected because of the OpenAI 2023-11-08 (94 min outage), Anthropic April 2026 (multi-hour Claude.ai/API/Code outage), and Anthropic Aug 2025 misroute (Sonnet 4 → 1M-context servers, 16% degraded) precedents — a multi-hour single-provider outage takes the entire platform down without multi-provider routing.

When it fails. Provider hard down → multi-provider failover; tenants whose tool-call format is provider-locked may fail mid-run (mitigation: keep normalized form in orchestrator, not provider-specific). Provider gray failure (5xx on a subset of model+region pairs while status page is green) → circuit-break per (provider, region, model) on observed error rate. Anthropic Aug/Sep 2025 postmortem documents three concurrent provider issues that mainstream multi-provider failover masked partially. OpenAI 2024-06-04 5+ hr ChatGPT outage is the precedent for hard-down.

External · MCP ServersAnthropic-spec MCP (OAuth 2.1 PKCE for remote)

Third-party MCP (Model Context Protocol) servers exposing capabilities to the agent: filesystem MCP, GitHub MCP, Linear MCP, Slack MCP, browser-use MCP, and per-tenant custom MCPs. Each server advertises a tool catalog via tools/list; the agent calls a specific tool via tools/call. Servers act as OAuth Resource Servers per the June 2025 MCP spec, advertising their auth endpoints via .well-known. Mediated entirely through mcp-router so the platform sees every tool call and can scope/filter/audit.

Why it exists. Considered first-party tool implementations only. Rejected because the MCP ecosystem hit 97 M monthly SDK downloads / 81 K GitHub stars by March 2026; building every integration in-house is non-leverage. The trade-off accepted: third-party MCP servers are an attack surface (Anthropic filesystem MCP CVE-2025-53109/53110 EscapeRoute, Anthropic Git MCP injection flaws fixed Jan 2026) — every MCP must be sandboxed by capability scoping and per-tool circuit breakers at mcp-router, NEVER directly exposed to the model.

When it fails. MCP server hard-down → per-tool circuit-breaker opens; agent sees 'tool unavailable' and pivots in the prompt. Compromised MCP server (supply-chain / CVE) → capability scoping limits blast radius; output classifier blocks exfil attempts. The Anthropic April 23 2026 'MCP applications on Claude.ai unavailable' (00:41-02:09 UTC, v2.1.120 force-rolled-back) is the precedent for fleet-wide MCP layer failure. EscapeRoute (CVE-2025-53109/53110) is the precedent for compromised-server-as-attack-vector.

Stage by stage

The same 10 stages the workspace walks, answered.

01Clarifications

What would you ask before drawing a single box?

Typical clarifications to surface:

  • What kind of agent? Coding (Devin, Cursor) — sandboxed shell + browser; web-research (Manus) — browser + summarization; ops/runbook (Lindy, internal SRE) — APIs only, no sandbox. The required tools change the topology.
  • Single-tenant or multi-tenant? Multi-tenant means per-tenant cost caps, sandbox isolation per task, capability tokens. Single-tenant is dramatically simpler.
  • Synchronous responses or background? Cursor inline = synchronous (10s budget). Devin background = asynchronous, hours-long. Different SSE / status APIs.
  • Self-hosted models or third-party? Anthropic + OpenAI is the default (frontier-quality matters for agent reliability). Self-hosted on vLLM is a vendor-lock-mitigation tier, not a default.
  • Compliance scope? SOC 2 + PCI is one tier; HIPAA / FedRAMP is another (sandboxed-data only, audit retention 7 y, BYOK at S3). Affects MCP allowlist and Object Lock posture.
  • Read-only or write-capable agents? A read-only research agent has no MCP write capabilities; a coding agent has filesystem + git + GitHub MCP write. Capability scoping is the line.

Assumptions to state:

  • 50K concurrent active runs (peak), mean run duration 15 min (Devin ACU = 15 min of active work; Devin Free 10 ACU/day, Pro 100/day).
  • 60 LLM calls / run × 12 K input tokens / call (heavy prompt with caching) → ~3 K LLM QPS at peak.
  • Anthropic Sonnet 4.6 default ($3/MTok input, $15/MTok output, 90% read discount on cached input).
  • Single home region with cross-region async DR follower; multi-region active-active is out of scope.
  • Sandbox: 2 vCPU / 1-2 GB / Firecracker microVM per task; warm pool for steady state, fork-from-snapshot for bursts.

02Functional reqs

What must this system actually do?

  • Start a run with an instruction, optional initial context, optional capability scope (allowlist of tools).
  • Stream the run's events via SSE: tokens, tool-call requests, tool results, status changes. Resumable via Last-Event-ID.
  • Inject a mid-run user message (pause-and-redirect pattern, like Cursor's "stop, do this instead").
  • Cancel a run at any time; orchestrator releases the sandbox, persists transcript, marks terminal.
  • Approve a HITL gate (run paused awaiting human approval — like Devin's "ready to merge?" prompt).
  • List / search runs by tenant, status, time, tags.
  • Retrieve a finished run's full transcript and result artifacts (code diffs, browser screenshots, generated files).
  • Replay a run (re-execute the same plan) for debugging.

03Non-functional

What must it promise about speed, uptime and correctness?

  • Durability: A run that has been ack'd by POST /v1/runs MUST complete (or terminate explicitly) — even across orchestrator restarts, sandbox-host loss, LLM-provider partial outages, and AZ-level failures.
  • Availability: 99.95% on POST /v1/runs (start). 99.9% on SSE event delivery — mid-run, the run survives even if SSE streams drop; the user reconnects with Last-Event-ID.
  • Latency: POST /v1/runs returns run_id + SSE URL in p99 < 500 ms. First token streamed in p99 < 2 s (Anthropic Sonnet 4.6 first-token p50 ~1.26 s per Artificial Analysis). Step-to-step latency p99 < 10 s.
  • Correctness under retry: Idempotent run-start (Brandur-style). Idempotent activities (Temporal default). Idempotent tool calls (per-tool idempotency_key forwarded through MCP). Idempotent mid-run signals — POST /v1/runs/:id/messages REQUIRES an Idempotency-Key (or client message_id); the workflow de-dupes by tracking seen message_ids in its state, so a retried injection (the transport retries 5xx) is applied at-most-once and never double-injected into agent context.
  • Cost containment: Per-tenant daily cap enforced pre-call at the gateway. No-progress detector kills loops in <60 s. Per-run hard cap on tokens.
  • Security: Firecracker microVM per task. OAuth 2.1 PKCE for MCP servers. Per-tool capability tokens. Output classifier on tool returns (prompt-injection defense). Audit log immutable for 7 y.

04Capacity estimation

How much load and data does this have to hold?

Defaults: 50K concurrent active runs (peak), 15 min mean run duration, 60 LLM calls / run, 40 tool calls / run, 12K input + 800 output tokens / call, 85% prompt-cache hit ratio.

  • LLM Gateway QPS: 50 K × 60 / (15 × 60) = 3 333 LLM QPS sustained, ~10 K peak (3× burst).
  • MCP / tool QPS: 50 K × 40 / 900 = 2 222 tool QPS sustained, ~7 K peak.
  • Monthly token volume: 50 K × 60 × 12 800 / 900 × 86 400 × 30 / 1e12 ≈ 110 trillion tokens / month — pro-rated to a smaller tenant base, this is per-platform. Anthropic publishes price; cost-worker enforces.
  • Cost / run uncached (Sonnet 4.6): 60 × (12 000 × $3/MTok + 800 × $15/MTok) / 1 M = $2.88 / run (≈ Anthropic's published Managed Agents example).
  • Cost / run with 85% cache hit: 60 × (12 000 × ($3 × 0.15 + $0.30 × 0.85) + 800 × $15) / 1 M = $1.23 / run — ~57% saving (input drops from $3.00 to an effective $0.705/MTok; output is uncached). This is why the gateway is cache-aware. Claude Code published 92% hit rate / 81% reduction.
  • Sandbox pool RAM: 50 K × 2 GB = 100 TB at peak. With 1.5× memory overcommit on 32 GB hosts (effective 48 GB / host of microVM packing) → 100 TB / 48 GB ≈ ~2 100 hosts. Fork-from-snapshot keeps cold-start <500 ms; warm-pool size = arrival rate × p95 cold-start × 1.5 (Little's Law on burst tier; 50 K runs / 900 s ≈ 55 sandboxes/sec sustained × ~2 s p95 fresh-boot cold-start → ~165 warm, more during bursts; E2B's ~10/sec figure is the fork-vs-warm cost crossover, not this platform's arrival rate).
  • Transcript storage (90 d): 3 333 events/sec × 86 400 × 90 × 4 KB ≈ 104 TB on S3 (≈ $2 400/mo at S3-IA). Glacier after 90 d.
  • SSE egress (peak): 50 K active streams × ~5 KB/s = 250 MB/s sustained. CDN absorbs.
  • Workflows in flight (Temporal): 50 K — well under Temporal's published 4 096-shard ceiling per cluster (each shard handles ~1 000 active workflows comfortably).

What flips at 10× scale (~500 K concurrent runs):

  • Temporal: single cluster ceiling reached; shard by tenant cell (one cluster per cell, multi-cluster replication).
  • Sandbox pool: ~32 K hosts; egress proxy concentration becomes a bottleneck → per-AZ MCP router replicas.
  • LLM Gateway: 30 K QPS peak — provider TPM/RPM limits hit; multi-key fan-out + enterprise contract negotiation.
  • Vector memory: pgvector saturates at ~10 M vectors / 100 QPS / tenant; switch the >P99 tenants to Turbopuffer or Qdrant.

05API design

What does the outside world call, and what comes back?

POST /v1/runs
Authorization: Bearer rk_live_<restricted>
Idempotency-Key: 5a8c1f2e-9b3d-4e6f-8a2b-1c3d4e5f6a7b
Content-Type: application/json

{
  "instruction": "Refactor the auth module to use OAuth 2.1 PKCE; open a PR.",
  "model": "claude-sonnet-4-6",
  "tools": ["filesystem", "github", "bash"],
  "capability_scope": "github:repo:acme/auth-svc:write",
  "tags": ["ci"],
  "budget_usd": 5.00,
  "timeout_seconds": 3600,
  "stream": true
}

201 Created
{
  "run_id": "run_3MQz8jXY",
  "status": "running",
  "events_url": "/v1/runs/run_3MQz8jXY/events",
  "created_at": "2026-05-09T12:34:56Z"
}

# Retry with same key, same body — replays cached response (idempotent)
201 Created (same response)

# Per-tenant kill-switch already engaged
402 Payment Required
{ "error": "tenant_budget_exceeded", "remaining_usd": 0 }

# Per-tenant concurrent-run cap hit
429 Too Many Requests
Retry-After: 30
GET /v1/runs/run_3MQz8jXY/events
Accept: text/event-stream
Last-Event-ID: 1234

200 OK
content-type: text/event-stream

id: 1235
event: token
data: {"text": "I'll start by reading the current auth module."}

id: 1236
event: tool_call
data: {"tool": "filesystem.read", "args": {"path": "/auth/oauth.go"}}

id: 1237
event: tool_result
data: {"tool": "filesystem.read", "result": "package auth\n..."}

id: 1238
event: status
data: {"status": "awaiting_human_approval", "reason": "ready to open PR"}

# Non-streaming long-poll mode (SSE-degradation fallback). With
# Accept: application/json + ?mode=poll&since=<event_id>, the SAME endpoint
# returns a bounded event batch + a cursor instead of an open stream. This is
# the shape the client SDK polls when the SSE edge is degraded.
GET /v1/runs/run_3MQz8jXY/events?mode=poll&since=1234&limit=200
Accept: application/json

200 OK
{
  "events": [
    { "id": 1235, "event": "token", "data": { "text": "I'll start by ..." } },
    { "id": 1236, "event": "tool_call", "data": { "tool": "filesystem.read" } }
  ],
  "next_since": 1238,
  "has_more": false
}
POST /v1/runs/run_3MQz8jXY/messages
Idempotency-Key: 9f2a4c7e-1b6d-4e3a-8c5f-2d7e9a1b3c4d
{ "message": "use a different OAuth library" }
202 Accepted
# Mid-run message injection is a non-idempotent state mutation; the
# Idempotency-Key (or a client message_id) is REQUIRED. The workflow tracks
# seen message_ids in its state and ignores a duplicate signal, so a transport
# retry of POST /messages does NOT double-inject the message into agent context.

POST /v1/runs/run_3MQz8jXY/cancel
202 Accepted

POST /v1/runs/run_3MQz8jXY/approve
202 Accepted
GET /v1/runs/run_3MQz8jXY
200 {
  "run_id": "run_3MQz8jXY",
  "status": "succeeded",
  "cost_usd": 1.23,
  "tokens": { "input": 720000, "output": 48000, "cached_input": 612000 },
  "transcript_url": "https://transcripts.s3/...?presigned=15min",
  "artifacts": [{ "type": "pr_url", "value": "https://github.com/acme/auth-svc/pull/421" }],
  "created_at": "2026-05-09T12:34:56Z",
  "terminal_at": "2026-05-09T12:49:12Z"
}
# List / search runs. Keyset (not offset) pagination on (created_at DESC, run_id)
# to ride the (tenant_id, status, created_at DESC) covering index — no deep-offset
# scan. next_cursor is opaque (a base64 of the last (created_at, run_id) pair);
# clients pass it back verbatim and never parse it.
GET /v1/runs?status=running&tag=ci&created_before=2026-05-09T00:00:00Z&limit=50&cursor=<opaque>
Authorization: Bearer rk_live_<restricted>

200 OK
{
  "runs": [
    { "run_id": "run_3MQz8jXY", "status": "running", "cost_usd": 0.42,
      "created_at": "2026-05-08T22:10:03Z" }
  ],
  "next_cursor": "eyJjcmVhdGVkX2F0IjoiMjAyNi0wNS0wOFQyMjoxMDowM1oiLCJydW5faWQiOiJydW5fM01Rejhq..."
}
# next_cursor is null when the page is the last one.
# Full transcript + result artifacts of a finished run (the audit/debug record).
GET /v1/runs/run_3MQz8jXY/transcript
200 { "run_id": "run_3MQz8jXY", "events_url": "...", "transcript_url": "https://transcripts.s3/...?presigned=15min" }

# Replay — re-execute the same plan against current code/prompt for debugging.
# Idempotency-Key REQUIRED; returns a NEW run_id linked to the source via replay_of.
POST /v1/runs/run_3MQz8jXY/replay
Idempotency-Key: b3e1d9a2-5c7f-4a8b-9e2d-6f1a3c5b7d9e
201 { "run_id": "run_7KpL2vWq", "replay_of": "run_3MQz8jXY", "status": "running" }

06Data model

What gets stored, and what is it looked up by?

runs (Postgres + Citus, sharded by tenant_id):

fieldtypenotes
run_idvarchar(24)opaque ID; PK (tenant_id, run_id) — Citus PKs must include the shard key
tenant_idbigintshard key
workflow_idvarchar(64)Temporal workflow_id (= run_id)
statusenumpending, running, awaiting_human, succeeded, failed, cancelled
budget_usdnumeric(10,2)user-set hard cap
cost_usdnumeric(10,4)running, updated by cost-worker
input_tokensbigintcumulative
output_tokensbigintcumulative
cached_input_tokensbigintcumulative (cache-hit visibility)
capability_scopetextcomma-separated MCP tool scopes
tagstext[]user labels from POST /v1/runs; GIN-indexed for the tag= filter
created_attimestamp
terminal_attimestamp?NULL while running

idempotency_keys (Brandur-style, same shard):

fieldtypenotes
keyvarchar(255)Idempotency-Key header; PK (tenant_id, key)
tenant_idbigintshard key + scope
request_fingerprintbyteahash(method, path, body)
recovery_pointenumstarted, run_persisted, workflow_signaled, finished
response_codeint?cached on completion
response_bodyjsonb?cached on completion
expires_attimestampnow() + 24 h

outbox (run-create relay; same shard as runs so the INSERT joins the SERIALIZABLE txn):

fieldtypenotes
outbox_idbigserialPK (tenant_id, outbox_id)
tenant_idbigintshard key
run_idvarchar(24)
topictextrun-create
payloadjsonbStartWorkflowExecution args (workflow_id = run_id)
created_attimestamp
relayed_attimestamp?NULL until the relay acks; rows reaped after

tenant_meter (cost-worker single-writer per tenant):

fieldtypenotes
tenant_idbigintPK
daily_budget_usdnumeric(10,2)product-tier driven
cost_today_usdnumeric(10,4)running, updated by cost-worker
runs_todayint
kill_switch_engaged_attimestamp?non-null = halt all new LLM calls
killswitch_reasontext?daily_budget_exceeded, runaway_loop, manual_ops

episodic_memory (pgvector, tenant-scoped):

fieldtypenotes
memory_iduuidPK (tenant_id, memory_id)
tenant_idbigintshard + scope
run_idvarchar(24)?nullable for tenant-global memories
texttextthe memory chunk
embeddingvector(1536)OpenAI text-embedding-3-large
created_attimestamp
recall_countintusage-driven retention

Indexes:

  • runs: PK (tenant_id, run_id); (tenant_id, status, created_at DESC) covering for the list/search endpoint; GIN on tags for the tag= filter (tag-filtered pages still keyset-paginate on (created_at DESC, run_id) after the GIN filter); partial on WHERE status IN ('running', 'awaiting_human') for the active-runs page.
  • idempotency_keys: PK (tenant_id, key); partial on WHERE expires_at < ... for the reaper.
  • outbox: PK (tenant_id, outbox_id); partial on WHERE relayed_at IS NULL for the relay's tail query.
  • tenant_meter: PK (tenant_id); partial on WHERE kill_switch_engaged_at IS NOT NULL for the ops dashboard.
  • episodic_memory: HNSW on embedding; (tenant_id, run_id) for run-scoped recall.

Database choice — recommended:

Postgres 16 + Citus (sharded by tenant_id) for runs, idempotency, tenant_meter, and pgvector — co-located so the run + idem CLAIM commits in one SERIALIZABLE transaction at run-svc. Cassandra RF=3 LOCAL_QUORUM for Temporal workflow history (Temporal's documented production default). Kafka RF=3 acks=all for the event bus. S3 + Object Lock for transcripts and audit (7-year compliance retention). Redis Cluster for the auth scope cache and the SSE status cache only — never source of truth.

Why not single Postgres for workflow history? At 50 K concurrent runs × ~500 events / run / hour = 25 M events/hour, Postgres' WAL fan-out becomes the bottleneck around 5-10 K writes/sec; Cassandra's leaderless quorum scales horizontally. Temporal supports both but only Cassandra at production scale.

Why not DynamoDB for the runs table? Works for idempotency (single-key conditional writes are perfect). Doesn't fit the runs index — list-by-tenant + filter-by-status + paginate-by-time is a SQL-shape query that DynamoDB models awkwardly.

Why not a dedicated vector DB by default? pgvector + pgvectorscale gives 471 QPS @ 99% recall on 50 M vectors (TigerData benchmark) — sufficient to ~10 M vectors / tenant, and co-locating memory with run state simplifies the operations story. Switch to Turbopuffer (cold-large) or Qdrant (low-latency hybrid) at the >10 M tier.

07High-level design

Which components handle a request, and in what order?

Architecture summary (17 components):

  1. CDN · SSE Edge (Cloudflare): TLS-terminate, WAF, absorb SSE long-lived TCP, per-IP rate limit.
  2. API Gateway (Envoy): per-tenant token bucket BEFORE forwarding, JWT validation, mTLS to upstream.
  3. Auth / Identity: OAuth 2.1 PKCE for users, hashed-key for service accounts, JWKS rotation with overlap window.
  4. Run Orchestrator (Write Plane): POST /v1/runs hits the Brandur-style idempotency state machine, commits run + idem in one txn, signals Temporal via outbox-relay.
  5. Stream Reader (Read Plane): SSE token + status events, resumable via Last-Event-ID, falls back from Kafka (24 h hot) to S3 transcript (cold).
  6. LLM Gateway: the central nervous system — multi-provider routing with KV-cache-aware sticky session affinity, Anthropic prompt caching with 1-h TTL header, per-tenant kill-switch enforcement (the load-bearing piece).
  7. MCP Tool Router: OAuth 2.1 PKCE, per-tool circuit breaker, capability scoping, output classifier (prompt-injection defense), cache_tools_list=true.
  8. Sandbox Pool (Firecracker): per-task microVMs forked from warm snapshots (~150 ms boot), unix-socket egress proxy, /workspace bound to per-run scratch volume, snapshot on idle.
  9. Temporal Cluster (Coordinator): durable workflow engine; workflow code is deterministic Go that drives the agent loop; LLM/tool/sandbox calls are activities with at-least-once + exponential backoff; signal-with-timeout for HITL.
  10. Workflow History (Cassandra): RF=3 LOCAL_QUORUM, sharded by workflow_id; the event-sourced source of truth.
  11. Run Metadata (Postgres + Citus + pgvector): runs index, idempotency, tenant_meter, episodic memory.
  12. Event Bus (Kafka, 256 partitions): cost meter (partition by tenant_id), SSE fan-out (partition by run_id), audit (RF=5).
  13. Cost Aggregator + Kill-Switch (Worker): consumes Kafka, single-writer per tenant, fires kill-switch via gRPC to llm-gw within 50 ms of breach.
  14. Transcripts + Snapshots (S3 + Object Lock): Pointer State payload offload, sandbox memory snapshots, audit WORM retention.
  15. External LLM Providers: Anthropic / OpenAI / Bedrock / Vertex; pinned to dated snapshots; circuit-broken on observed error rate (NOT status page).
  16. External MCP Servers: OAuth 2.1 Resource Servers; mediated through mcp-router with output classification.
  17. (Implicit: monitoring + Langfuse/OTel sink consumes the audit topic.)

> SSE-degradation fallback (read-plane note): the documented client fallback when the CDN/SSE edge degrades is the long-poll mode of GET /v1/runs/:id/events?mode=poll&since=<event_id> (defined in # api) — NOT GET /v1/runs/:id?since=, which is only a terminal-status snapshot and carries no event-delta semantics. The poll path rides the existing CDN → gateway → stream-svc read edge; no new edge is required. (The CDN node's failureMode string still names the old ?since= shape and should be re-pointed at the events poll endpoint.) > List / transcript / replay endpoints ride existing edges: GET /v1/runs and GET /v1/runs/:id/transcript are reads served by stream-svc (CDN → gateway → stream-svc); POST /v1/runs/:id/replay is a write served by run-svc (gateway → run-svc) that starts a fresh workflow. No new HLD node/edge is needed for them.

Data flow on POST /v1/runs (synchronous slice): Client → CDN (WAF + SSE pre-ack) → Gateway (auth + per-tenant token bucket + signed scope envelope) → run-svc (Brandur idem CLAIM in SERIALIZABLE Postgres txn that also INSERTs the run row + an outbox row) → outbox-relay drains to Temporal StartWorkflowExecution → return 201 with run_id + SSE URL. <500 ms p99.

Data flow inside a workflow step (the hot path): Temporal worker picks up the next workflow task → workflow code decides "call LLM with current context" → schedules an LLM activity → activity dispatches to llm-gw → llm-gw checks tenant_meter (kill-switch?), picks provider via cache-aware sticky route, calls Anthropic with prompt-cache headers → streams tokens back (token deltas coalesced into ~4 KB / ~250 ms chunk events published to event-bus partitioned by run_id; stream-svc fans out via SSE) → activity completes with the model response → workflow code parses tool_use blocks → schedules a tool activity → activity dispatches to mcp-router → mcp-router validates capability token, runs the per-tool circuit breaker, calls the MCP server → result returns through output classifier → activity completes with the tool result → workflow code loops.

Data flow on sandbox crash mid-step: Sandbox host dies. Temporal's history service detects the missed activity heartbeat (Heartbeat timeout fires); the attempt fails and the activity retry policy reschedules it on a new worker — ScheduleToClose remains the overall cap across retries. New sandbox is acquired (warm pool or fork-from-snapshot), fetches the last sandbox snapshot from transcripts/S3, replays the activity. Workflow code is unaware — same activity_id, same input, same idempotency-keyed result.

Data flow on cost-runaway detection: llm-gw publishes llm.call event (tokens, cost) to Kafka per call. cost-worker consumes (single-writer per tenant via Kafka sticky partition), updates tenant_meter, evaluates thresholds. On breach: writes kill_switch_engaged_at = now(), gRPC to llm-gw → llm-gw rejects all subsequent calls for that tenant with 402 Payment Required. Total time from breaching call to kill-switch enforcement: <50 ms p99.

08Deep dives

Which part breaks first, and what do you do about it?

1. Why an agent run is a durable workflow. The defining decision in the architecture. Three properties only durable workflows give you: (a) Survives orchestrator restart — Temporal replays from history; the workflow picks up exactly where it left off, on a different machine if needed. (b) Survives sandbox host loss — the workflow rebinds the sandbox activity to a new microVM; the long agent transcript is in S3 (Pointer State offload) so reload is cheap. (c) Survives LLM provider partial outage — the LLM activity retries with exponential backoff; if the per-(provider, region, model) circuit breaker opens, llm-gw fails over to the next-best provider; workflow code keeps a normalized tool-call schema so failover is possible mid-run. The trade-off: workflow code MUST be deterministic. All non-determinism (LLM responses, tool results, time, randomness) lives in activities. Adding Workflow.GetVersion gates around every prompt-change deploy is non-negotiable; without it, the dreaded temporal_workflow_task_failed{cause="non_deterministic_error"} pages oncall on every shipped change.

2. KV-cache-aware sticky routing. Manus's published lesson: KV-cache hit rate is the single most important metric for production agents. Anthropic's prompt-caching read price is 10% of the base input price (90% discount); a 90% cache hit rate cuts input cost by ~80%. The mechanism: vLLM's prefix caching and SGLang's RadixAttention hash prompt prefixes; if the next request to the same upstream pod has the same prefix, the KV cache is reused. The gateway requirement: sticky session affinity by hash(tenant_id, prompt_prefix_hash) so each replica owns ~1/N of the prefix-hash space and pins follow-up requests to the upstream pod that already holds the cache. Round-robin defeats this. Prompt-construction discipline matters too: Manus's lesson — append-only context, deterministic JSON serialization, stable system-prompt prefix; a single token change at the start of the prefix invalidates the cache. Anthropic-specific: set the anthropic-beta: extended-cache-ttl-2025-04-11 header and cache_control: { ttl: "1h" } because Anthropic silently dropped Claude Code's default TTL from 1 h to 5 min in April 2026 (The Register documented this). Claude Code achieves 92% cache hit / 81% cost reduction; Manus reports 80%; we target 85%.

3. Cost-runaway kill-switch — alerting is not enforcement. The published $47K LangChain incident: four agents in an unintended A2A loop ran for 11 days, $47K bill. The retro: token-budget alerts are not budget enforcement. The architecture: every llm-gw call publishes (tenant_id, run_id, tokens, cost) to Kafka; cost-worker consumes (single-writer per tenant via Kafka sticky partition), aggregates into tenant_meter, and on threshold breach writes kill_switch_engaged_at + sends a gRPC kill-switch to llm-gw. llm-gw's pre-call check rejects all subsequent calls for that tenant with 402. Three thresholds: (a) daily budget — hard cap from billing. (b) per-minute runaway — Free $0.50/min, Pro $5/min — catches loops mid-incident. (c) no-progress detector — 5 identical (tool_name, args_hash) within 60 s = halt; this catches the LangGraph #6731 / LangChain #26019 infinite-tool-loop pattern. End-to-end latency from breaching call to kill-switch enforcement: <50 ms p99.

4. Sandbox isolation — Firecracker vs gVisor vs runc. Cognition's published "Building Cloud Agents" notes: "containerized agents share a kernel, and a single compromised session can reach every other container's filesystem." runc Leaky Vessels (CVE-2024-21626) and the Nov 2025 trio (CVE-2025-31133/52565/52881) prove container-class isolation is too thin for agent-generated code. Firecracker (KVM hardware virtualization) is the production primitive every coding-agent vendor (E2B, Cognition, Vercel Sandbox, Devin) converged on — 125 ms boot, <5 MB RAM overhead per microVM, 4 000 demonstrated per host. Modal uses gVisor (user-space kernel intercepting syscalls); chosen for cold-start economics, weaker isolation than Firecracker. The trade-off accepted in this canonical: Firecracker for security, with hybrid pool sizing — always-on warm pool for steady-state arrival (sub-150 ms acquire), fork-from-snapshot for bursts (still <500 ms). Even so, the Anthropic filesystem MCP CVE-2025-53109/53110 (EscapeRoute, path-traversal) and Anthropic's April 2025 Devin-Sliver-download incident prove sandboxing ≠ MCP safety; capability scoping at mcp-router and an output classifier on tool returns are defense-in-depth.

5. Indirect prompt injection — the lethal trifecta. Simon Willison's "lethal trifecta" framing: a deployment that combines (private data + untrusted content + external comms) is a prompt-exfiltration vulnerability waiting to happen. EchoLeak (CVE-2025-32711, June 2025): zero-click M365 Copilot exfil via crafted email; CVSS 9.3. Devin April 2025: a poisoned GitHub issue caused Devin to fetch an attacker site, download a Sliver C2 binary, chmod +x it, and execute (sysid blog: "your agent has root"). Three load-bearing mitigations in this canonical: (a) Capability scoping at mcp-router — every tool call is scoped to (run, tool, OAuth scope) at issue time; the model cannot use a tool not in its scope. (b) Egress allowlist at the sandbox — the unix-socket egress proxy blocks any destination not in the per-tenant allowlist; the agent literally cannot reach an attacker site even if it tries. (c) Output classifier on tool returns — IndirectInjectionClassifier blocks instruction-shaped tool output before it reaches the next LLM call. The two-LLM pattern (planner sees no untrusted text) is referenced in the deep-dive prompt but layered on top, not load-bearing for the canonical.

6. Resume-on-crash protocol. A 2-hour Devin task. Sandbox host dies at minute 90. What survives, why, and how cheaply does the run resume? What survives: the Temporal workflow history in Cassandra (event-sourced; every plan-step + activity-result is logged). The transcript chunks + sandbox memory snapshots in S3 (Pointer State offload — Temporal history holds only the S3 URL, not the 100 MB transcript). The kill-switch state in tenant_meter. What's lost: the in-flight microVM's RAM-only state (any work since the last snapshot). The protocol: Temporal's history service detects the missed heartbeat on the sandbox activity (Heartbeat timeout fires); the attempt fails and the retry policy reschedules the activity on a new worker — ScheduleToClose remains the overall cap across retries. New sandbox is acquired (warm or forked); fetches the last snapshot from S3 (~500 ms load); workflow code resumes from the next event. Total resume cost: ~2 s of overhead. The load-bearing piece: snapshot frequency. Snapshot every 30 s during long-running steps means worst-case 30 s of lost work; snapshot every 5 min means 5 min — pick the trade per your run-length distribution. Devin's published architecture takes the 30-s-snapshot path.

7. Hot tenant on a viral run — single-shard meltdown. A single tenant suddenly spawns 1 000 concurrent runs (Cursor user reports of this exist on HN). Tenant's Citus shard (run-meta) saturates; tenant_meter row sees write contention from cost-worker; Kafka partition for that tenant_id sees consumer lag; kill-switch decision latency climbs from 50 ms to 5 s. Mitigations: (a) per-tenant concurrent-run cap at the gateway (Manus uses 2/3/10 by tier — never let one tenant exceed their slice). (b) tenant_meter shard hot-row protection — SELECT … FOR UPDATE SKIP LOCKED + 1-second batched updates rather than per-event writes. (c) cost-worker partition assignment strategy: cooperative-sticky avoids rebalance storms during scale events. The Stripe 2019-07-10 single-shard gray failure retro is the cautionary lesson — health checks pass, correctness invariants break. SLO: kill-switch decision latency p99 under sustained 1 000-runs/tenant load.

8. Workflow versioning and the non_deterministic_error pager. The most insidious failure mode in a Temporal-backed agent platform. You ship a new agent prompt. Workflows from yesterday were authored against the old prompt. They replay. The agent loop's control-flow now diverges (different tool selected at step 3 because the new prompt biases differently). Temporal raises temporal_workflow_task_failed{cause="non_deterministic_error"} and the workflow halts mid-run. The fix at code time: every prompt-changing deploy gates via Workflow.GetVersion("agent-prompt-v2", 0, 1); new logic only runs when the workflow was started after the version bump. Replay tests in CI via workflowcheck catch missed gates. The on-call playbook: non_deterministic_error rate >0 is a P1 page; rollback the deploy; re-run affected workflows from a known-good version. Temporal's community has multi-year threads on this exact failure mode.

09Trade-offs

What did this design cost, and what breaks at 10×?

What breaks at 10× scale (~500K concurrent runs):

  • Temporal cluster: 4 096-shard ceiling reached; shard by tenant cell (one cluster per cell, multi-cluster replication). Cell-based architecture per AWS / Stripe pattern.
  • Sandbox pool: ~32 K hosts; egress proxy concentration becomes a bottleneck → per-AZ MCP router replicas with a stateless egress proxy in each. Also: warm-pool size grows 10× — pre-warming becomes its own cost line.
  • LLM Gateway: 30 K LLM QPS at peak — provider TPM/RPM limits hit on a single API key; multi-key fan-out + enterprise contract negotiation.
  • Cassandra (workflow history): 256 shards comfortably; at 10× consider expanding to 1 024 + multi-DC replication. ScyllaDB is the lower-latency drop-in; Discord's published Cassandra→ScyllaDB migration is the precedent.
  • Vector memory: pgvector saturates around 10 M vectors / 100 QPS / tenant; switch the >P99 tenants to Turbopuffer (cold-large, $50/M/mo) or Qdrant (low-latency hybrid).
  • Kafka: 256 partitions stretched; expand to 512 on agent.events.v1; consider tiered storage (Confluent / WarpStream) for the hot tier.

Per-component failure stories:

  • Idempotency store down: run-svc fails CLOSED 503 (cannot accept run start without a CLAIM). In-region sync replication = same story as Stripe's 2019 ledger.
  • Temporal frontend down: run-svc StartWorkflowExecution returns 503; outbox-relay retries with exponential backoff; clients see brief 202s with status URLs. Workflow history is unaffected.
  • Workflow history (Cassandra) quorum loss: workflows on affected shards stall (read+write block). Monzo 2019-07-29 retro is the lesson — never relax to ONE on writes.
  • LLM provider hard down: circuit-break, multi-provider failover. Tool-call format normalized in orchestrator (NOT provider-specific) so failover is possible mid-run.
  • Sandbox host loss: activity reschedules; fresh microVM loads last snapshot; ~2 s of overhead. Worst-case lost work bounded by snapshot interval.
  • MCP server outage: per-tool circuit opens; agent sees 'tool unavailable' and pivots; Anthropic April 23 2026 MCP outage is the precedent.
  • Cost-worker lag: kill-switch fires late; tenant overspends. Mitigation: gateway hard-cap at 1.5× tenant budget bounds the loss even with full cost-worker failure.
  • Region failure (us-east-1): Temporal multi-cluster replication promotes standby (manual; 10 min RTO). GitHub 2018-10-21 (43-second partition split-brain MySQL) is the lesson on automatic vs manual cross-region failover.
  • Bad model deploy (Anthropic Aug 25 2025 XLA bug): snapshot-pinned model references mean a regressed alias doesn't propagate; canary new snapshots on shadow traffic; rollback gate keyed on tool_call_parse_error_rate.
  • Bad agent prompt deploy: Workflow.GetVersion gates prevent non-determinism on in-flight runs; replay tests in CI catch missed gates. non_deterministic_error_rate > 0 is the P1 page.

Trade-offs we accepted:

  • Single home region: active-active multi-region needs Spanner-class storage for both run-meta and Cassandra history; we picked the simpler Postgres+Cassandra stack with async cross-region DR (manual promote, ~10 min RTO; data-loss budget = 0 because reconciliation closes on promotion).
  • Frontier-only models: self-hosted on vLLM is the vendor-lock-mitigation tier, not the default. Frontier model quality (Claude Opus 4.7 / GPT-5) is the load-bearing piece for agent reliability. Open-weight self-hosted is the contingency, not the steady state.
  • Snapshot every 30 s during long-running activities: worst-case 30 s of lost work on sandbox crash. Tunable per run-class; we picked the published Devin-style number.
  • Idempotency window 24 h: beyond 24 h, a same-key replay is treated as a new run. Bug-fix headroom 48 h soft-retain (Brandur). Misconfigured clients retrying past that window is on them.
  • HITL pause auto-resolves at 4 h: prevents the resource-leak deadlock pattern documented in Temporal community. Tunable per run; the canonical takes the safe default.
  • MCP servers as third-party attack surface: capability scoping at mcp-router + sandbox egress allowlist + output classifier are the defense-in-depth. Some compromises survive these (output classifier false-negative rate is published 5-15%); the residual risk is accepted with audit-log + per-tenant blast-radius bounding.

Primary sources

  • Cognition — Devin's 2025 Performance Review
  • Cognition — What We Learned Building Cloud Agents
  • Manus — Context Engineering for AI Agents (KV-cache hit rate)
  • Anthropic — Prompt Caching docs (5min/1h TTL)
  • Anthropic — Postmortem of Three Recent Issues (Aug/Sep 2025)
  • Anthropic — Claude Code sandboxing
  • Temporal — Replit Agent case study
  • Temporal — Of course you can build dynamic AI agents
  • Diagrid — Checkpoints Are Not Durable Execution
  • E2B — Firecracker vs QEMU (125 ms boot, 4 000 microVMs/host)
  • Modal — Top AI Code Sandbox Products 2025
  • Simon Willison — The lethal trifecta for AI agents
  • OWASP — Top 10 for Agentic Applications (Dec 2025)
  • Anthropic — EscapeRoute MCP CVEs (CVE-2025-53109/53110)
  • Replit — Effort-based pricing recap (July 2025 billing bug)
  • OpenTelemetry — GenAI semantic conventions (gen_ai.operation.name)

Now defend it

Reading a design is not the same as being able to hold one under questioning. The workspace asks the same questions an interviewer would, and the simulator disagrees with you when the diagram does not support the claim.

Work AI Agent Platform yourself

Build the primitives this design leans on

Each one is an animated curriculum that constructs the system from scratch.

More in Consensus, Coordination & Durable Execution

Getting N machines to agree, and getting one job to happen exactly once: Raft, coordination services, CRDTs, locks, leader election, schedulers and durable workflows.