✳flâneur — a map of the web's best reading
Language Model Circuits Are Sparse in the Neuron Basis | Transluce AI
transluce.org · 18,651 words · saved by 1 readers
We introduce a new approach to finding sparse circuits in the neuron basis, without relying on learned features.
Language Model Circuits Are Sparse in the Neuron Basis Aryaman Arora * , Zhengxuan Wu * , Jacob Steinhardt , Sarah Schwettmann * Equal contribution. Correspondence to: aryaman@transluce.org, zen@transluce.org. Transluce | Published: November 20, 2025 Many interpretability methods rely on learned feature bases—such as sparse autoencoders or cross-layer transcoders—based on the belief that neurons do not cleanly decompose model computation. We revisit this assumption and show that, with a better choice of neuron basis (MLP activations) and a stronger attribution method (RelP), raw neurons can pr
Explore this link on the map →related reading
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWronglesswrong.com
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Toy Models of Superpositiontransformer-circuits.pub
- Transluce on X: "Is your LM secretly an SAE? Most circuit-finding interpretability methods use learned features rather than raw activations, based on the belief that neurons do not cleanly decompose computation. In our new work, we show MLPx.com
- Zoom In: An Introduction to Circuitsdistill.pub
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net
- Sparse Attention Post-Training for Mechanistic Interpretabilityarxiv.org