[2406.04093] Scaling and evaluating sparse autoencoders
Abstract:Sparse autoencoders provide a promising unsupervised approach for extracting interpretable features from a language model by reconstructing activations from a sparse bottleneck layer. Since language models learn many concepts, autoencoders need to be very large to recover all relevant features. However, studying the properties of autoencoder scaling is difficult due to the need to balance reconstruction and sparsity objectives and the presence of dead latents. We propose using k-sparse autoencoders [Makhzani and Frey, 2013] to directly control sparsity, simplifying tuning and improving the reconstruction-sparsity frontier. Additionally, we find modifications that result in few dead latents, even at the largest scales we tried. Using these techniques, we find clean scaling laws with respect to autoencoder size and sparsity. We also introduce several new metrics for evaluating feature quality based on the recovery of hypothesized features, the explainability of activation patterns, and the sparsity of downstream effects. These metrics all generally improve with autoencoder size. To demonstrate the scalability of our approach, we train a 16 million latent autoencoder on GPT-4 activations for 40 billion tokens. We release training code and autoencoders for open-source models, as well as a visualizer.
[2406.04093] Scaling and evaluating sparse autoencoders Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Machine Learning arXiv:2406.04093 (cs) [Submitted on 6 Jun 2024] Title: Scaling and evaluating sparse autoencoders Authors: Leo Gao , Tom Dupré la Tour , Henk Tillman , Gabriel Goh , Rajan Troll , Alec Radford , Ilya Sutskever , Jan Leike , Jeffrey Wu View a PDF of the paper titled Scaling and evaluating sparse autoencoders, by Leo Gao and 8 other authors View PDF HTML (experimen
Explore this link on the map →related reading
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- [Interim research report] Taking features out of superposition with sparse autoencoders — LessWronglesswrong.com
- A gentle introduction to sparse autoencodersnickjiang.substack.com
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- pdfopenreview.net
- Open Source Automated Interpretability for Sparse Autoencoder Features | EleutherAI Blogblog.eleuther.ai
- [2605.29358] Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnetarxiv.org
- Matryoshka Sparse Autoencoders — LessWronglesswrong.com
- An Intuitive Explanation of Sparse Autoencoders for Mechanistic Interpretability of LLMs — LessWronglesswrong.com