flâneur — a map of the web's best reading

Softmax Linear Units

transformer-circuits.pub · 8,966 words · saved by 2 readers

As Transformer generative models continue to gain real-world adoption , it becomes ever more important to ensure they behave predictably and safely, in both the short and long run.  Mechanistic interpretability – the project of attempting to reverse engineer neural networks into understandable computer programs – offers one possible avenue for addressing these safety issues: by understanding the internal structures that cause neural networks to produce the outputs they do, it may be possible to address current safety problems more systematically as well as anticipating future safety problems. Until recently mechanistic interpretability has focused primarily on CNN vision models , but some recent efforts have begun to explore mechanistic interpretability for transformer language models .  Notably, we were able to reverse-engineer 1 and 2 layer attention-only transformers and we used empirical evidence to draw indirect conclusions about in-context learning in arbitrarily large models .

Softmax Linear Units Transformer Circuits Thread Softmax Linear Units Authors Nelson Elhage ∗† , Tristan Hume ∗ , Catherine Olsson ∗ , Neel Nanda ∗§ , Tom Henighan † , Scott Johnston † , Sheer El Showk † , Nicholas Joseph † , Nova DasSarma † , Ben Mann † , Danny Hernandez, Amanda Askell, Kamal Ndousse, Andy Jones, Dawn Drain, Anna Chen, Yuntao Bai, Deep Ganguli, Liane Lovitt, Zac Hatfield-Dodds , Jackson Kernion, Tom Conerly, Shauna Kravec, Stanislav Fort, Saurav Kadavath, Josh Jacobson, Eli Tran-Johnson, Jared Kaplan, Jack Clark, Tom Brown, Sam McCandlish, Dario Amodei, Christopher Olah ‡ Aff

Explore this link on the map →

saved by

related reading