Don't use cosine similarity carelessly
p.migdal.pl · 3,131 words · saved by 1 readers
Cosine similarity - the duct tape of AI. Convenient but often misused. Let's find out how to use it better.
Don't use cosine similarity carelessly 14 Jan 2025 | by Piotr Migdał see Hacker News for the discussion Midas turned everything he touched into gold. Data scientists turn everything into vectors. We do it for a reason — as gold is the language of merchants, vectors are the language of AI. Just as Midas discovered that turning everything to gold wasn’t always helpful, we’ll see that blindly applying cosine similarity to vectors can lead us astray. While embeddings do capture similarities, they often reflect the wrong kind - matching questions to questions rather than questions to answers, or ge
related reading
- Vector Similarity Explained | Pineconepinecone.io
- Comparison of different Word Embeddings on Text Similarity — A use case in NLP | by Intellica.AI | Mediumintellica-ai.medium.com
- Embeddings: What they are and why they mattersimonwillison.net
- The Illustrated Word2vec – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Sentence Embeddings. Introduction to Sentence Embeddings – hackerllamaosanseviero.github.io
- Vector embeddings | OpenAI APIdevelopers.openai.com
- Word Embeddingslena-voita.github.io
- Vector embeddings | OpenAI APIplatform.openai.com
- Semantic Word Embeddingsoffconvex.org
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Introducing text and code embeddings | OpenAIopenai.com
- Announcing ScaNN: Efficient Vector Similarity Searchai.googleblog.com