flâneur — a map of the web's best reading

Interpreting Language Model Parameters

goodfire.ai · 27,108 words · saved by 4 readers

Neural networks use millions to trillions of parameters to learn how to solve tasks that no other machines can solve. What structure do these parameters learn? And how do they compute intelligent behavior? Mechanistic interpretability aims to uncover how neural networks use their parameters to implement their impressive neural algorithms. Although previous work has uncovered substantial structure in the intermediate representations that networks use, little progress has been made to understand how the parameters and nonlinearities of networks perform computations on those representations. In this work, we present a method that brings us closer to this understanding by decomposing a language model's parameters into subcomponents that each implement only a small part of the model's learned algorithm, while simultaneously requiring only a small fraction of those subcomponents to account for the network's behavior on any input. The method, adVersarial Parameter Decomposition (VPD), optimiz

Interpreting Language Model Parameters Research Interpreting Language Model Parameters Authors Lucius Bushnaq 1,* Dan Braun 1,*,† Oliver Clive-Griffin 1,*,† Bart Bussmann 2 Nathan Hu 2 Michael Ivanitskiy 2 Linda Linsefors 3 Lee Sharkey 1,* 1 Goodfire 2 MATS 3 Independent * Core contributor. † Equal contribution; order randomized. Correspondence to lee@goodfire.com See also our Contributions Statement . Find the markdown version of this post here . Published May 5th 2026 Neural networks use millions to trillions of parameters to learn how to solve tasks that no other machines can solve. What st

Explore this link on the map →

saved by

related reading