✳flâneur — a map of the web's best reading
Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWrong
lesswrong.com · 4,259 words · saved by 3 readers
TLDR: Recently, Gao et al trained transformers with sparse weights, and introduced a pruning algorithm to extract circuits that explain performance o…
x Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWrong Interpretability (ML & AI) AI Frontpage 2026 Top Fifty: 9 % 136 Weight-Sparse Circuits May Be Interpretable Yet Unfaithful by jacob_drori 9th Feb 2026 9 min read 5 136 TLDR: Recently, Gao et al trained transformers with sparse weights, and introduced a pruning algorithm to extract circuits that explain performance on narrow tasks. I replicate their main results and present evidence suggesting that these circuits are unfaithful to the model’s “true computations”. This work was done as part of the Anthropic Fellows Program
Explore this link on the map →saved by
related reading
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Attribution-based parameter decomposition — LessWronglesswrong.com
- Transformer Circuits Threadtransformer-circuits.pub
- Sparse Attention Post-Training for Mechanistic Interpretabilityarxiv.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Attribution Patching: Activation Patching At Industrial Scale - Neel Nandaneelnanda.io
- Language Model Circuits Are Sparse in the Neuron Basis | Transluce AItransluce.org
- Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research] — LessWronglesswrong.com