06
Web Crawler
Politely traverse the web at scale. Don't crawl yourself in circles.SavedSaved on this device — Saved on this device
01Clarifications
What would you ask before drawing a single box?
Ambiguity you would resolve with the interviewer: scope, scale, who uses it, what counts as done.
AI staff engineer
Enter to send · Shift+Enter for a new line
About Web Crawler
Politely traverse the web at scale. Don't crawl yourself in circles.
- Difficulty
- advanced
- Time
- about 50 minutes
- Stages
- 10
- Topic
- Search, Indexing & Retrieval
How this problem is worked
Ten stages, from the questions you would ask an interviewer to the trade-offs you would defend. Each asks one question, and the simulator runs the architecture you draw against the requirements you wrote.
- 01ClarificationsWhat would you ask before drawing a single box?
- 02Functional reqsWhat must this system actually do?
- 03Non-functionalWhat must it promise about speed, uptime and correctness?
- 04Capacity estimationHow much load and data does this have to hold?
- 05API designWhat does the outside world call, and what comes back?
- 06Data modelWhat gets stored, and what is it looked up by?
- 07Use-case breakdownHow does each requirement actually get served?
- 08High-level designWhich components handle a request, and in what order?
- 09Deep divesWhich part breaks first, and what do you do about it?
- 10Trade-offsWhat did this design cost, and what breaks at 10×?
Primary sources for this problem
- Mercator: A Scalable, Extensible Web Crawler — Heydon & Najork (Compaq SRC, 1999)
- Detecting Near-Duplicates for Web Crawling — Manku, Jain, Das Sarma (Google, WWW 2007)
- Inside Googlebot — Google Search Central Blog (Mar 2026)
- Heritrix Politeness Parameters — Internet Archive Wiki
- Common Crawl monthly statistics & WARC file format
- Cloudflare AI Audit — blocking AI crawlers with one click
- Cloudflare on Perplexity stealth crawlers (Aug 2025)
- RFC 9309 — Robots Exclusion Protocol
- Google open-source robots.txt parser (google/robotstxt)
- Apache StormCrawler vs Nutch (DZone)
More in Search, Indexing & Retrieval
Finding a needle: inverted indexes, distributed search, vector and ANN retrieval, crawling, and the query box itself.
- Build Build a distributed search engine (Elasticsearch / OpenSearch style)Five million books, a search box, and a 100 ms budget. Build the engine from the inverted index up — segment, refresh, shard, replica, scatter-gather, BM25 — and feel why every guarantee that lives across shards is paid for in either an extra round trip or a small lie about the rankings.
- Build Build a vector database (Pinecone / Weaviate / pgvector style)Approximate nearest-neighbor over a billion 1536-dim vectors in 10 ms. Build the index from scratch (HNSW, IVF, PQ), pay the recall-vs-latency tax explicitly, support filtered + hybrid search, and feel why every LLM stack in 2026 has a vector store next to its KV store.
Browse the full problem catalog, or see what the simulator does and does not model.