Towards Scalable Parameter Decomposition
The most successful methods so far, like SAEs, have focused largely on the activations that flow through a model, rather than directly inspecting the weights that transform and guide the inputs into those flows. It's a bit like trying to understand a program by only looking at its runtime variables, but never its source code. Parameter decomposition offers a way to decompose a model’s parameters—the 'source code'—into components that reveal not only what the network computes, but how it computes it. Today, we're releasing a paper on Stochastic Parameter Decomposition (SPD), which removes key barriers to the scalability of prior methods.
Towards Scalable Parameter Decomposition Research Towards Scalable Parameter Decomposition Authors Lucius Bushnaq *† Dan Braun *† Lee Sharkey ‡ * Co-first Author † Goodfire – Work primarily carried out while at Apollo Research ‡ Goodfire Correspondence to lucius@goodfire.ai Blog post by Michael Byun Published June 27, 2025 Full Paper Read on arXiv → Like a game of Jenga, SPD finds parameter components that can be removed while keeping the model's behavior stable. "Jenga" by Sheila Sund, licensed under CC BY 2.0 . In 2013, researchers discovered that word2vec embeddings encoded analogies in the
Explore this link on the map →related reading
- Towards Scalable Parameter Decompositiongoodfire.ai
- Stochastic Parameter Decompositionarxiv.org
- Interpreting Language Model Parametersgoodfire.ai
- Attribution-based parameter decomposition — LessWronglesswrong.com
- [2506.20790] Stochastic Parameter Decompositionarxiv.org
- [2501.14926] Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decompositionarxiv.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Activation space interpretability may be doomed — LessWronglesswrong.com
- Transformer Circuits Threadtransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org