Infrastructure Primitives
Design a Web Crawler
A distributed crawler that discovers, fetches, and indexes billions of pages while respecting robots.txt, avoiding duplicate work, and controlling politeness per host.
Open and Simulate this Architecture in InfraDraftCore Architectural Components
URL Frontier
Priority queue balancing crawl-worthiness against per-host politeness delay.
Distributed Fetcher Workers
Horizontally-scaled workers pulling URLs and downloading page content.
DNS Resolution & Cache
Avoids repeated DNS lookups from becoming the crawl bottleneck.
Robots.txt Parser & Politeness Scheduler
Enforces crawl-delay and disallow rules per domain.
Duplicate/Near-Duplicate Content Detector
Content hashing or SimHash to skip re-processing unchanged or mirrored pages.
Link Extraction & Normalization
Parses outbound links and canonicalizes URLs before re-queueing.
Distributed Crawl Store
Blob storage for raw fetched HTML, feeding downstream indexing.
Seen-URL Bloom Filter
Memory-efficient probabilistic check to avoid re-crawling known URLs.