flâneur — a map of the web's best reading

How the Residual Stream is (Not) Linear

cs.columbia.edu · saved by 1 readers

I commonly hear about the linearity of the residual stream of Transformer language models. Linearity, it is argued, is powerful for interpretability, and gives validity to interpretability tools like logit lens and steering vectors. In this brief note, we ask, linear with respect to what? We argue that the residual stream is not linear with respect to anything that gives us significant leverage. The additive nature of updates to the residual stream is still interesting, but not as powerful as linearity. Futher, empirically, assuming additive changes correspond to generalizing behavior change seems to work in some cases, perhaps because of what functions are easy to learn with residual connections. However, I encourage researchers not to refer to the Transformer's architecture to assign interpretability/control value to methods that implicitly assume linearity, and to instead let those methods' empirical results stand alone. If you're not familiar with these concepts, I encourage you to

I commonly hear about the linearity of the residual stream of Transformer language models. Linearity, it is argued, is powerful for interpretability, and gives validity to interpretability tools like logit lens and steering vectors. In this brief note, we ask, linear with respect to what? We argue that the residual stream is not linear with respect to anything that gives us significant leverage. The additive nature of updates to the residual stream is still interesting, but not as powerful as linearity. Futher, empirically, assuming additive changes correspond to generalizing behavior change s

Explore this link on the map →