flâneur — a map of the web's best reading

Sparse mixtures of linear transforms

transformer-circuits.pub · 2,894 words · saved by 1 readers

This work is a preliminary update describing a method we are still developing. As such, many of the results we present here are incomplete. We present them in the hopes that it inspires follow-up work and improvements from the external research community. Upon publishing this update, we became aware of contemporaneous work on a highly-related architecture (“Mixture of Decoders,” Oldfield et al.). We have updated the post to describe the similarities / differences between our methods in the Related Work section. We recommend that anyone interested in this work read theirs as well! In our recent work, we trained transcoders – sparse, extra-wide MLPs – as more interpretable replacements for a model’s original MLP layers. We used the transcoder neurons (“features”) as a basis for understanding model computation. We described computations using “attribution graphs,” which depict the causal interactions between features that give rise to the model’s outputs. While this approach has proved v

Sparse mixtures of linear transforms Transformer Circuits Thread Sparse mixtures of linear transforms Jack Lindsey, Brian Chen, Adam Pearce, Sasha Hydrie, Thomas Conerly; edited by Jeff Wu This work is a preliminary update describing a method we are still developing. As such, many of the results we present here are incomplete. We present them in the hopes that it inspires follow-up work and improvements from the external research community. Upon publishing this update, we became aware of contemporaneous work on a highly-related architecture (“Mixture of Decoders,” Oldfield et al. ). We have up

Explore this link on the map →

related reading