ICLR Poster Demystifying Embedding Spaces using Large Language Models
Embeddings have become a pivotal means to represent complex, multi-faceted information about entities, concepts, and relationships in a condensed and useful format. Nevertheless, they often preclude direct interpretation. While downstream tasks make use of these compressed representations, meaningful interpretation usually requires visualization using dimensionality reduction or specialized machine learning interpretability methods. This paper addresses the challenge of making such embeddings more interpretable and broadly useful, by employing large language models (LLMs) to directly interact with embeddings -- transforming abstract vectors into understandable narratives. By injecting embeddings into LLMs, we enable querying and exploration of complex embedding data. We demonstrate our approach on a variety of diverse tasks, including: enhancing concept activation vectors (CAVs), communicating novel embedded entities, and decoding user preferences in recommender systems. Our work coupl
ICLR Poster Demystifying Embedding Spaces using Large Language Models ICLR 2024 CSP Test --> Poster Demystifying Embedding Spaces using Large Language Models Guy Tennenholtz ⋅ Yinlam Chow ⋅ ChihWei Hsu ⋅ Jihwan Jeong ⋅ Lior Shani ⋅ Azamat Tulepbergenov ⋅ Deepak Ramachandran ⋅ Martin Mladenov ⋅ Craig Boutilier 2024 Poster [ OpenReview ] Abstract Embeddings have become a pivotal means to represent complex, multi-faceted information about entities, concepts, and relationships in a condensed and useful format. Nevertheless, they often preclude direct interpretation. While downstream tasks make use
Explore this link on the map →related reading
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub
- Embeddings: What they are and why they mattersimonwillison.net
- Transformer Circuits Threadtransformer-circuits.pub
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Vector embeddings | OpenAI APIdevelopers.openai.com
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- GitHub - PaulPauls/llama3_interpretability_sae: A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible. · GitHubgithub.com
- Vector embeddings | OpenAI APIplatform.openai.com
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Emotion Concepts and their Function in a Large Language Modeltransformer-circuits.pub
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub