2501.14926
arxiv.org · 8,138 words · saved by 1 readers
N/A
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition Dan Braun∗ Lucius Bushnaq∗ Stefan Heimersheim∗ Jake Mendel arXiv:2501.14926v4 [cs.LG] 7 Feb 2025 Lee Sharkey† Apollo Research‡…
saved by
related reading
- Towards Scalable Parameter Decompositiongoodfire.ai
- Stochastic Parameter Decompositionarxiv.org
- Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decompositionarxiv.org
- [2501.14926] Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decompositionarxiv.org
- Attribution-based parameter decomposition — LessWronglesswrong.com
- Interpreting Language Model Parametersgoodfire.ai
- Stochastic Parameter Decompositionarxiv.org
- [2506.20790] Stochastic Parameter Decompositionarxiv.org
- Towards Scalable Parameter Decompositiongoodfire.ai
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Toy Models of Superpositiontransformer-circuits.pub