← InfraDraft

Infrastructure Primitives

Design a Web Crawler

A distributed crawler that discovers, fetches, and indexes billions of pages while respecting robots.txt, avoiding duplicate work, and controlling politeness per host.

Open and Simulate this Architecture in InfraDraft

Core Architectural Components

URL Frontier

Priority queue balancing crawl-worthiness against per-host politeness delay.

Distributed Fetcher Workers

Horizontally-scaled workers pulling URLs and downloading page content.

DNS Resolution & Cache

Avoids repeated DNS lookups from becoming the crawl bottleneck.

Robots.txt Parser & Politeness Scheduler

Enforces crawl-delay and disallow rules per domain.

Duplicate/Near-Duplicate Content Detector

Content hashing or SimHash to skip re-processing unchanged or mirrored pages.

Link Extraction & Normalization

Parses outbound links and canonicalizes URLs before re-queueing.

Distributed Crawl Store

Blob storage for raw fetched HTML, feeding downstream indexing.

Seen-URL Bloom Filter

Memory-efficient probabilistic check to avoid re-crawling known URLs.