Transluce on X: "Is your LM secretly an SAE? Most circuit-finding interpretability methods use learned features rather than raw activations, based on the belief that neurons do not cleanly decompose computation. In our new work, we show MLP neurons actually do support sparse, faithful circuits! https://t.co/lTBbUqoRlt" / Twitter
To view keyboard shortcuts, press question mark View keyboard shortcuts Added to your Bookmarks Add to Folder Tweet See new posts Conversation Neil Rathi reposted Transluce @TransluceAI Is your LM secretly an SAE? Most circuit-finding interpretability methods use learned features rather than raw activations, based on the belief that neurons do not cleanly decompose computation. In our new work, we show MLP neurons actually do support sparse, faithful circuits! 11:00 AM · Nov 20, 2025 ·Twitter Web App 9 Quote Tweets 5 74 249 216 Tweet your reply Reply Transluce @TransluceAI · Nov 20 Comparing with sparse autoencoders (SAEs), we show that circuits traced directly on a model’s MLP neurons can be just as sparse and faithful. We use two advances to achieve this result. 1 22 Transluce @TransluceAI · Nov 20 First, we use MLP neurons instead of MLP outputs as the units of our circuits. Due to the preceding nonlinearity, features are encouraged to align with the MLP neuron basis. 1 21 Transl
Transluce @TransluceAI Is your LM secretly an SAE? Most circuit-finding interpretability methods use learned features rather than raw activations, based on the belief that neurons do not cleanly decompose computation. In our new work, we show MLP neurons actually do support sparse, faithful circuits! 7:00 PM · Nov 20, 2025 133.2K Views 7 0 7 77 0 7 7 376 0 3 7 6 345 0 3 4 5 Read 7 replies
Explore this link on the map →related reading
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Toy Models of Superpositiontransformer-circuits.pub
- Language Model Circuits Are Sparse in the Neuron Basis | Transluce AItransluce.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Transformer Circuits Threadtransformer-circuits.pub
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Zoom In: An Introduction to Circuitsdistill.pub
- Activation space interpretability may be doomed — LessWronglesswrong.com
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- A List of 45+ Mech Interp Project Ideas from Apollo Research’s Interpretability Team — AI Alignment Forumalignmentforum.org