book3n.dvi
infolab.stanford.edu · 10,194 words · saved by 1 readers
N/A
72 Chapter 3 Finding Similar Items A fundamental data-mining problem is to examine data for “similar” items. We shall take up applications in Section 3.1, but an example would be looking at a collection of Web pages and finding near-duplicate pages. These pages could be plagiarisms, for example, or they could be mirrors that have almost the same content but differ in information about the host and about other mirrors. The naive approach to finding pairs of similar items requires us to look at ev- ery pair of items. When we are dealing with a large dataset, looking at all pairs of…
related reading
- Introduction to Locality-Sensitive Hashingtylerneylon.com
- Faiss: A library for efficient similarity search - Engineering at Metaengineering.fb.com
- Announcing ScaNN: Efficient Vector Similarity Searchai.googleblog.com
- Hierarchical Navigable Small Worlds (HNSW) | Pineconepinecone.io
- GitHub - facebookresearch/faiss: A library for efficient similarity search and clustering of dense vectors.github.com
- How does Audio Fingerprinting work - Emysoundemysound.com
- [2011.10427] Dataset Discovery in Data Lakesarxiv.org
- Bloom Filters - Much, much more than a space efficient hashmap! | Ben E. C. Boyterboyter.org
- GitHub - nullnull/simstring: A Python implementation of the SimString, a simple and efficient algorithm for approximate string matching. · GitHubgithub.com
- Fast regex search: indexing text for agent tools · Cursorcursor.com
- HyperLogLog - Wikipediaen.wikipedia.org
- Hierarchical Navigable Small Worlds (HNSW) | Pineconepinecone.io