Sparse mixtures of linear transforms
This work is a preliminary update describing a method we are still developing. As such, many of the results we present here are incomplete. We present them in the hopes that it inspires follow-up work and improvements from the external research community. Upon publishing this update, we became aware of contemporaneous work on a highly-related architecture (“Mixture of Decoders,” Oldfield et al.). We have updated the post to describe the similarities / differences between our methods in the Related Work section. We recommend that anyone interested in this work read theirs as well! In our recent work, we trained transcoders – sparse, extra-wide MLPs – as more interpretable replacements for a model’s original MLP layers. We used the transcoder neurons (“features”) as a basis for understanding model computation. We described computations using “attribution graphs,” which depict the causal interactions between features that give rise to the model’s outputs. While this approach has proved v
Sparse mixtures of linear transforms Transformer Circuits Thread Sparse mixtures of linear transforms Jack Lindsey, Brian Chen, Adam Pearce, Sasha Hydrie, Thomas Conerly; edited by Jeff Wu This work is a preliminary update describing a method we are still developing. As such, many of the results we present here are incomplete. We present them in the hopes that it inspires follow-up work and improvements from the external research community. Upon publishing this update, we became aware of contemporaneous work on a highly-related architecture (“Mixture of Decoders,” Oldfield et al. ). We have up
Explore this link on the map →related reading
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Transformers from Scratche2eml.school
- [2406.11944] Transcoders Find Interpretable LLM Feature Circuitsarxiv.org
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Tracing Attention Computation Through Feature Interactionstransformer-circuits.pub
- Softmax Linear Unitstransformer-circuits.pub
- A List of 45+ Mech Interp Project Ideas from Apollo Research’s Interpretability Team — AI Alignment Forumalignmentforum.org
- Monet: Mixture of Monosemantic Experts for Transformers Explained — LessWronglesswrong.com