✳flâneur — a map of the web's best reading
Building a web search engine from scratch in two months with 3 billion neural embeddings
blog.wilsonl.in · 8,712 words · saved by 14 readers
End-to-end deep dive of the project, spanning a large GPU cluster, distributed RocksDB, and terabytes of sharded HNSW.
A while back, I decided to undertake a project to challenge myself: build a web search engine from scratch. Aside from the fun deep dive opportunity, there were two motivators: Search engines seemed to be getting worse, with more SEO spam and less relevant quality content. Transformer-based text embedding models were taking off and showing amazing natural comprehension of language. A simple question I had was: why couldn't a search engine always result in top quality content? Such content may be rare, but the Internet's tail is long , and better quality results should rank higher than the prol
Explore this link on the map →saved by
- Amir
- Elizabeth Qiu
- Rajan Agarwal
- Aryan Naik
- Karthik Suresh
- anka hu
- Jianmin Chen
- Daniel Kiss
- Shreyas Prasad
- Siddharth Ramakrishnan
- Rohan Kanti
- Nicklaus Tran
related reading
- The Anatomy of a Search Engineinfolab.stanford.edu
- Not All Vector Databases Are Made Equal | Towards Data Sciencetowardsdatascience.com
- Hierarchical Navigable Small Worlds (HNSW) | Pineconepinecone.io
- How We Store and Search 30 Billion Facesclearview.ai
- Crawling a billion web pages in just over 24 hoursandrewkchan.dev
- Our AI Research: How We Evaluate Semantic Search Technology | Exa Blogexa.ai
- Embeddings: What they are and why they mattersimonwillison.net
- Faiss: A library for efficient similarity search - Engineering at Metaengineering.fb.com
- Announcing ScaNN: Efficient Vector Similarity Searchai.googleblog.com
- Vector embeddings | OpenAI APIdevelopers.openai.com
- Perfect Web Search for AI Agents with Semantic Search Technology | Exa Blogexa.ai
- turbopuffer: fast search on object storageturbopuffer.com