Toy Models of Superposition
It would be very convenient if the individual neurons of artificial neural networks corresponded to cleanly interpretable features of the input. For example, in an “ideal” ImageNet classifier, each neuron would fire only in the presence of a specific visual feature, such as the color red, a left-facing curve, or a dog snout. Empirically, in models we have studied, some of the neurons do cleanly map to features. But it isn't always the case that features correspond so cleanly to neurons, especially in large language models where it actually seems rare for neurons to correspond to clean features. This brings up many questions. Why is it that neurons sometimes align with features and sometimes don't? Why do some models and tasks have many of these clean neurons, while they're vanishingly rare in others? In this paper, we use toy models — small ReLU networks trained on synthetic data with sparse input features — to investigate how and when models represent more features than they have dime
Toy Models of Superposition Transformer Circuits Thread Toy Models of Superposition Authors Nelson Elhage ∗ , Tristan Hume ∗ , Catherine Olsson ∗ , Nicholas Schiefer ∗ , Tom Henighan , Shauna Kravec, Zac Hatfield-Dodds , Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg ∗ , Christopher Olah ‡ Affiliations Anthropic, Harvard Published Sept 14, 2022 * Core Research Contributor; ‡ Correspondence to colah@anthropic.com ; Author contributions statement below . It would be very convenient if the individual neurons of artificial neural
saved by
- Winnie Xu
- Katherine Huang
- Alex K. Chen
- Tazik Sh
- Claire Wang
- Alex Shaw
- Lowell Dennings
- Ulisse mini
- Hangyul Lyna Kim
- Braeden Hall
- Katherine He
- Lydia Nottingham
related reading
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- What Would Non-Linear Features Actually Look Like? — Liv Gortonlivgorton.com
- Softmax Linear Unitstransformer-circuits.pub
- Interpretability Dreamstransformer-circuits.pub
- how neural networks think at scalemarmik.xyz
- The Building Blocks of Interpretabilitydistill.pub
- [2608.27540] Towards a mathematical theory of superpositionarxiv.org
- Fact Finding: Attempting to Reverse-Engineer Factual Recall on the Neuron Level (Post 1) — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io