✳flâneur — a map of the web's best reading
Attribution-based parameter decomposition — LessWrong
lesswrong.com · 4,357 words · saved by 1 readers
This is a linkpost for Apollo Research's new interpretability paper: …
x Attribution-based parameter decomposition — LessWrong Interpretability (ML & AI) Apollo Research (org) AI Frontpage 2025 Top Fifty: 14 % 109 Attribution-based parameter decomposition by Lucius Bushnaq , Dan Braun , StefanHex , jake_mendel , Lee Sharkey 25th Jan 2025 AI Alignment Forum Linkpost for publications.apolloresearch.ai 5 min read 21 109 Ω 45 This is a linkpost for Apollo Research's new interpretability paper: " Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition ". We introduce a new method for directly decomp
Explore this link on the map →saved by
related reading
- Stochastic Parameter Decompositionarxiv.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Interpreting Language Model Parametersgoodfire.ai
- AlgZoo: uninterpreted models with fewer than 1,500 parameters — LessWronglesswrong.com
- Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWronglesswrong.com
- Towards Scalable Parameter Decompositiongoodfire.ai
- [2501.14926] Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decompositionarxiv.org
- Towards Scalable Parameter Decompositiongoodfire.ai
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- [2506.20790] Stochastic Parameter Decompositionarxiv.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org