Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWrong
lesswrong.com · 4,259 words · saved by 3 readers
TLDR: Recently, Gao et al trained transformers with sparse weights, and introduced a pruning algorithm to extract circuits that explain performance o…
x Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWrong Interpretability (ML & AI) AI Frontpage 2026 Top Fifty: 9 % 136 Weight-Sparse Circuits May Be Interpretable Yet Unfaithful by jacob_drori 9th Feb 2026 9 min read 5 136 TLDR: Recently, Gao et al trained transformers with sparse weights, and introduced a pruning algorithm to extract circuits that explain performance on narrow tasks. I replicate their main results and present evidence suggesting that these circuits are unfaithful to the model’s “true computations”. This work was done as part of the Anthropic Fellows Program
saved by
related reading
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Transformer Circuits Threadtransformer-circuits.pub
- Attribution-based parameter decomposition — LessWronglesswrong.com
- Sparse Attention Post-Training for Mechanistic Interpretabilityarxiv.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Towards Automated Circuit Discovery for Mechanistic Interpretabilityarxiv.org
- Attribution Patching: Activation Patching At Industrial Scale - Neel Nandaneelnanda.io
- Language Model Circuits Are Sparse in the Neuron Basis | Transluce AItransluce.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub