Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
Using a sparse autoencoder, we extract a large number of interpretable features from a one-layer transformer. Browse A/1 Features → Browse All Features → Not published yet. No DOI yet. Mechanistic interpretability seeks to understand neural networks by breaking them into components that are more easily understood than the whole. By understanding the function of each component, and how they interact, we hope to be able to reason about the behavior of the entire network. The first step in that program is to identify the correct components to analyze. Unfortunately, the most natural computational unit of the neural network – the neuron itself – turns out not to be a natural unit for human understanding. This is because many neurons are polysemantic: they respond to mixtures of seemingly unrelated inputs. In the vision model Inception v1, a single neuron responds to faces of cats and fronts of cars . In a small language model we discuss in this paper, a single neuron responds to a mixture
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning Transformer Circuits Thread Towards Monosemanticity: Decomposing Language Models With Dictionary Learning Towards Monosemanticity: Decomposing Language Models With Dictionary Learning Using a sparse autoencoder, we extract a large number of interpretable features from a one-layer transformer. Browse A/1 Features → Browse All Features → Authors Trenton Bricken * , Adly Templeton * , Joshua Batson * , Brian Chen * , Adam Jermyn * , Tom Conerly, Nicholas L Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby,
Explore this link on the map →saved by
- Winnie Xu
- Claire Wang
- HudZah
- Rishi Kothari
- Asher P
- Nima Pourjafar
- Vyom Pathak
- Aaron Pham
- Dhruv Gautam
- Sidney Nimako
- Lydia Nottingham
- Marc Esmael Karimi
related reading
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Toy Models of Superpositiontransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- Transformer Circuits Threadtransformer-circuits.pub
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- [2309.08600] Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Circuits Updates - January 2024transformer-circuits.pub