flâneur — a map of the web's best reading

Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWrong

lesswrong.com · 4,259 words · saved by 3 readers

TLDR: Recently, Gao et al trained transformers with sparse weights, and introduced a pruning algorithm to extract circuits that explain performance o…

x Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWrong Interpretability (ML & AI) AI Frontpage 2026 Top Fifty: 9 % 136 Weight-Sparse Circuits May Be Interpretable Yet Unfaithful by jacob_drori 9th Feb 2026 9 min read 5 136 TLDR: Recently, Gao et al trained transformers with sparse weights, and introduced a pruning algorithm to extract circuits that explain performance on narrow tasks. I replicate their main results and present evidence suggesting that these circuits are unfaithful to the model’s “true computations”. This work was done as part of the Anthropic Fellows Program

Explore this link on the map →

saved by

related reading