Crawling a billion web pages in just over 24 hours
For some reason, nobody’s written about what it takes to crawl a big chunk of the web in a while: the last point of reference I saw was Michael Nielsen’s post from 2012[1]. Obviously lots of things have changed since then. Most bigger, better, faster: CPUs have gotten a lot more cores, spinning disks have been replaced by NVMe solid state drives with near-RAM I/O bandwidth, network pipe widths have exploded, EC2 has gone from a tasting menu of instance types to a whole rolodex’s worth, yada yada. But some harder: much more of the web is dynamic, with heavier content too. How has the state of the art changed? Have the bottlenecks shifted, and would it still cost ~$41k to bootstrap your own Google? I wanted to find out, so I built and ran my own web crawler1 under similar constraints. Time limit of 24 hours. Because I thought a billion pages crawled in a day was achievable based on preliminary experiments and 40 hours doesn’t sound as cool. In my final crawl, the average active time of e
Contents Crawling a billion web pages in just over 24 hours, in 2025 Discussion on r/programming . tl;dr: 1.005 billion web pages 25.5 hours $462 For some reason, nobody's written about what it takes to crawl a big chunk of the web in a while: the last point of reference I saw was Michael Nielsen's post from 2012 . Obviously lots of things have changed since then. Most bigger, better, faster: CPUs have gotten a lot more cores, spinning disks have been replaced by NVMe solid state drives with near-RAM I/O bandwidth, network pipe widths have exploded, EC2 has gone from a tasting menu of instance
Explore this link on the map →saved by
related reading
- Building a web search engine from scratch in two months with 3 billion neural embeddingsblog.wilsonl.in
- The Anatomy of a Search Engineinfolab.stanford.edu
- The Architecture of a Web Crawler: Building a Google-Inspired Distributed Web Crawler. Part 1 | by TonyWang | Mediummedium.com
- Production Twitter on One Machine? 100Gbps NICs and NVMe are fast - Tristan Humethume.ca
- This Page is Designed to Last: A Manifesto for Preserving Content on the Webjeffhuang.com
- Populating the page: how browsers work - Performance | MDNdeveloper.mozilla.org
- Rediscovering the Small Web - Neustadt.frneustadt.fr
- Why your website should be under 14kB in size | endtimes.devendtimes.dev
- How web bloat impacts users with slow connectionsdanluu.com
- Speed Up Web Scraping Using Concurrency and Parallelismscrapehero.com
- BigPipe: Pipelining web pages for high performance - Engineering at Metaengineering.fb.com
- Aggressive AI scrapers are making it kinda suck to run wikis | Weird Gloopweirdgloop.org