Sparse Autoencoders Reveal Universal Feature Spaces Across Large Language Models
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. We investigate feature universality in large language models (LLMs), a research field that aims to understand how different models similarly represent concepts in the latent spaces of their intermediate layers. Demonstrating feature universality allows discoveries about latent representations to generalize across several models. However, comparing features across LLMs is challenging due to polysemanticity, in which individual neurons often correspond to multiple features rather than distinct ones. This makes it difficult to disentangle and match features across different models. To address this issue, we employ a method known as dictionary learning by using sparse autoencoders (SAEs) to transform LLM activations into more interpretable spaces spanned by
Sparse Autoencoders Reveal Universal Feature Spaces Across Large Language Models Michael Lan Philip Torr ‡ Austin Meek § Ashkan Khakzar ‡ David Krueger ♡ Fazl Barez †‡ † Tangentic ‡ University of Oxford § University of Delaware ♡ MILA Work done during the ERA-Krueger AI Safety Lab internship. Author contributions detailed in § Author Contributions . Abstract We investigate feature universality in large language models (LLMs), a research field that aims to understand how different models similarly represent concepts in the latent spaces of their intermediate layers. Demonstrating feature univer
Explore this link on the map →related reading
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- pdfopenreview.net
- A gentle introduction to sparse autoencodersnickjiang.substack.com
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- Matryoshka Sparse Autoencoders — LessWronglesswrong.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- [2409.14507] A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencodersar5iv.labs.arxiv.org
- Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Modelsarxiv.org
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com