flâneur — a map of the web's best reading

Understanding the Parameter Decomposition papers

logangraves.com · 13 words · saved by 1 readers

Before reading this post, I recommend reading Lee Sharkey's post "Mech Interp is not Pre-paradigmatic." I think the post is worthwhile in itself as a survey of mech interp methods and history; it also explains Parameter Decomposition's basic ideas. Ideally you will have familiarity with those before reading this post, because I don't spend a ton of time outlining them. This is mostly an explanation of the technical details behind what Goodfire and Sharkey are shooting for with Parameter Decomposition. This post closely follows the content of the APD paper Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition and the SPD paper Stochastic Parameter Decomposition. I follow the notation there, but I sometimes use my own structure. To recap, PD takes all the weight matrices of the model and flattens them into a vector in 'parameter-space'. Then it aims to turn this vector into the sum of a bunch of other, simpler weight

404 Error 404 Not Found There's nothing at this URL :( (try searching?)

Explore this link on the map →

related reading