Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
Eight months ago, we demonstrated that sparse autoencoders could recover monosemantic features from a small one-layer transformer. At the time, a major concern was that this method might not scale feasibly to state-of-the-art transformers and, as a result, be unable to practically contribute to AI safety. Since then, scaling sparse autoencoders has been a major priority of the Anthropic interpretability team, and we're pleased to report extracting high-quality features from Claude 3 Sonnet,Anthropic's medium-sized production model. We find a diversity of highly abstract features. They both respond to and behaviorally cause abstract behaviors. Examples of features we find include features for famous people, features for countries and cities, and features tracking type signatures in code. Many features are multilingual (responding to the same concept across languages) and multimodal (responding to the same concept in both text and images), as well as encompassing both abstract and concre
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet Transformer Circuits Thread Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet Authors Adly Templeton * , Tom Conerly * , Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, Alex Tamkin, Esin Durmus, Tristan Hume, Francesco Mosconi, C. Daniel Freeman, Theodore R. Sumers, Ed
Explore this link on the map →saved by
- Winnie Xu
- Alicia Guo
- Elizabeth Qiu
- Karan MJ
- Claire Wang
- Emma Guo
- Benedict Neo
- Elliot Kim
- Vyom Pathak
- Katherine Driscoll
- Alex Becker
- Lydia Nottingham
related reading
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- [2605.29358] Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnetarxiv.org
- Transformer Circuits Threadtransformer-circuits.pub
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Neuronpedianeuronpedia.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io