[2406.01506] The Geometry of Categorical and Hierarchical Concepts in Large Language Models
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
View PDF HTML (experimental) Abstract:The linear representation hypothesis is the informal idea that semantic concepts are encoded as linear directions in the representation spaces of large language models (LLMs). Previous work has shown how to make this notion precise for representing binary concepts that have natural contrasts (e.g., {male, female}) as directions in representation space. However, many natural concepts do not have natural contrasts (e.g., whether the output is about an animal). In this work, we show how to extend the formalization of the linear representation hypothesis to…
saved by
related reading
- The Linear Representation Hypothesis and the Geometry of Large Language Modelsarxiv.org
- On the Origins of Linear Representations in Large Language Modelsarxiv.org
- The Linear Representation Hypothesis and the Geometry of Large Language Modelsarxiv.org
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Emotion Concepts and their Function in a Large Language Modeltransformer-circuits.pub
- [2412.06769] Training Large Language Models to Reason in a Continuous Latent Spacearxiv.org
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Large Concept Models (LCMs) by Meta: The Era of AI After LLMs?aipapersacademy.com
- [2602.15029] Symmetry in language statistics shapes the geometry of model representationsarxiv.org
- Training Large Language Models to Reason in a Continuous Latent Spacearxiv.org
- [2310.06824] The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasetsarxiv.org