[2409.14507] A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
Sparse Autoencoders (SAEs) have emerged as a promising approach to decompose the activations of Large Language Models (LLMs) into human-interpretable latents. In this paper, we pose two questions. First, to what extent…
A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders David Chanin 1,2,* , James Wilken-Smith 1,* , Tomáš Dulka 1,* , Hardik Bhatnagar 1,* , Joseph Bloom 1,3 1 LASR Labs, 2 University College London, 3 Decode Research * These authors contributed equally to this work Abstract Sparse Autoencoders (SAEs) have emerged as a promising approach to decompose the activations of Large Language Models (LLMs) into human-interpretable latents. In this paper, we pose two questions. First, to what extent do SAEs extract monosemantic and interpretable latents? Second, to what e
Explore this link on the map →related reading
- pdfopenreview.net
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Matryoshka Sparse Autoencoders — LessWronglesswrong.com
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- A gentle introduction to sparse autoencodersnickjiang.substack.com
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- An Intuitive Explanation of Sparse Autoencoders for Mechanistic Interpretability of LLMs — LessWronglesswrong.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- I Trained a Language Model. Then I Built a Brain Scanner and Looked Inside It. | by Caleb DeLeeuw | Mediummedium.com
- Do sparse autoencoders find "true features"? — LessWronglesswrong.com
- Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning — LessWronglesswrong.com