AI Agent Platform
Long-running multi-step LLM agents — durable workflow + sandboxed execution + LLM gateway. Brain / hands / state are independently replaceable. The agent run is a workflow, not a request.01Clarifications
What would you ask before drawing a single box?
Ambiguity you would resolve with the interviewer: scope, scale, who uses it, what counts as done.
Enter to send · Shift+Enter for a new line
About AI Agent Platform
Long-running multi-step LLM agents — durable workflow + sandboxed execution + LLM gateway. Brain / hands / state are independently replaceable. The agent run is a workflow, not a request.
- Difficulty
- advanced
- Time
- about 50 minutes
- Stages
- 10
- Topic
- Consensus, Coordination & Durable Execution
How this problem is worked
Ten stages, from the questions you would ask an interviewer to the trade-offs you would defend. Each asks one question, and the simulator runs the architecture you draw against the requirements you wrote.
- 01ClarificationsWhat would you ask before drawing a single box?
- 02Functional reqsWhat must this system actually do?
- 03Non-functionalWhat must it promise about speed, uptime and correctness?
- 04Capacity estimationHow much load and data does this have to hold?
- 05API designWhat does the outside world call, and what comes back?
- 06Data modelWhat gets stored, and what is it looked up by?
- 07Use-case breakdownHow does each requirement actually get served?
- 08High-level designWhich components handle a request, and in what order?
- 09Deep divesWhich part breaks first, and what do you do about it?
- 10Trade-offsWhat did this design cost, and what breaks at 10×?
Primary sources for this problem
- Cognition — Devin's 2025 Performance Review
- Cognition — What We Learned Building Cloud Agents
- Manus — Context Engineering for AI Agents (KV-cache hit rate)
- Anthropic — Prompt Caching docs (5min/1h TTL)
- Anthropic — Postmortem of Three Recent Issues (Aug/Sep 2025)
- Anthropic — Claude Code sandboxing
- Temporal — Replit Agent case study
- Temporal — Of course you can build dynamic AI agents
- Diagrid — Checkpoints Are Not Durable Execution
- E2B — Firecracker vs QEMU (125 ms boot, 4 000 microVMs/host)
- Modal — Top AI Code Sandbox Products 2025
- Simon Willison — The lethal trifecta for AI agents
- OWASP — Top 10 for Agentic Applications (Dec 2025)
- Anthropic — EscapeRoute MCP CVEs (CVE-2025-53109/53110)
- Replit — Effort-based pricing recap (July 2025 billing bug)
- OpenTelemetry — GenAI semantic conventions (gen_ai.operation.name)
Build the primitives this design leans on
Each one is an animated curriculum that constructs the system from scratch.
More in Consensus, Coordination & Durable Execution
Getting N machines to agree, and getting one job to happen exactly once: Raft, coordination services, CRDTs, locks, leader election, schedulers and durable workflows.
- Build Build Raft — consensus you can defendReplicate a deterministic state machine across N servers with safety as a theorem and liveness under partial synchrony. Build the protocol from term to commit to safety proof to reads, and feel why etcd, Cockroach, and TiKV ship slightly different Rafts.
- Build Build a workflow engine (Temporal / Airflow / Cadence style)A function that survives crashes, restarts, and re-deploys — and still finishes. Build a durable execution engine where workflow code is replayed deterministically from an event history, activities retry with exponential backoff, sagas compensate on failure, and the same workflow definition runs identically a year later. Internalize why 'just retry the cron job' breaks at the second step.
- Distributed LockRedlock controversy, fencing tokens, lease vs lock.
Browse the full problem catalog, or see what the simulator does and does not model.