Web Crawler

What would you ask before drawing a single box?

Ambiguity you would resolve with the interviewer: scope, scale, who uses it, what counts as done.

About Web Crawler

Politely traverse the web at scale. Don't crawl yourself in circles.

Difficulty
advanced
Time
about 50 minutes
Stages
10
Topic
Search, Indexing & Retrieval

How this problem is worked

Ten stages, from the questions you would ask an interviewer to the trade-offs you would defend. Each asks one question, and the simulator runs the architecture you draw against the requirements you wrote.

  1. 01ClarificationsWhat would you ask before drawing a single box?
  2. 02Functional reqsWhat must this system actually do?
  3. 03Non-functionalWhat must it promise about speed, uptime and correctness?
  4. 04Capacity estimationHow much load and data does this have to hold?
  5. 05API designWhat does the outside world call, and what comes back?
  6. 06Data modelWhat gets stored, and what is it looked up by?
  7. 07Use-case breakdownHow does each requirement actually get served?
  8. 08High-level designWhich components handle a request, and in what order?
  9. 09Deep divesWhich part breaks first, and what do you do about it?
  10. 10Trade-offsWhat did this design cost, and what breaks at 10×?

Primary sources for this problem

  • Mercator: A Scalable, Extensible Web Crawler — Heydon & Najork (Compaq SRC, 1999)
  • Detecting Near-Duplicates for Web Crawling — Manku, Jain, Das Sarma (Google, WWW 2007)
  • Inside Googlebot — Google Search Central Blog (Mar 2026)
  • Heritrix Politeness Parameters — Internet Archive Wiki
  • Common Crawl monthly statistics & WARC file format
  • Cloudflare AI Audit — blocking AI crawlers with one click
  • Cloudflare on Perplexity stealth crawlers (Aug 2025)
  • RFC 9309 — Robots Exclusion Protocol
  • Google open-source robots.txt parser (google/robotstxt)
  • Apache StormCrawler vs Nutch (DZone)

Browse the full problem catalog, or see what the simulator does and does not model.