In-context Learning and Induction Heads
As Transformer generative models continue to scale and gain increasing real world use , addressing their associated safety problems becomes increasingly important. Mechanistic interpretability – attempting to reverse engineer the detailed computations performed by the model – offers one possible avenue for addressing these safety issues. If we can understand the internal structures that cause Transformer models to produce the outputs they do, then we may be able to address current safety problems more systematically, as well as anticipating safety problems in future more powerful models. In the past, mechanistic interpretability has largely focused on CNN vision models, but recently, we presented some very preliminary progress on mechanistic interpretability for Transformer language models. Specifically, in our prior work we developed a mathematical framework for decomposing the operations of transformers, which allowed us to make sense of small (1 and 2 layer attention-only) models
In-context Learning and Induction Heads Transformer Circuits Thread In-context Learning and Induction Heads Authors Catherine Olsson ∗ , Nelson Elhage ∗ , Neel Nanda ∗ , Nicholas Joseph † , Nova DasSarma † , Tom Henighan † , Ben Mann † , Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds , Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, Chris Olah ‡ Affiliation Anthropic Published Mar 8, 2022 * Core Research Contributor; † Core Infrastructur
saved by
- Winnie Xu
- Amir
- Varun Shenoy
- Claire Wang
- Asma Lamgh
- Emma Guo
- Derek Yen
- Lydia Nottingham
- Timothy Kostolansky
- Julian H
- Akira Yoshiyama
- Samuel Chen
related reading
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Induction heads - illustrated — LessWronglesswrong.com
- [2212.07677] Transformers learn in-context by gradient descentarxiv.org
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Softmax Linear Unitstransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Mechanistic Interpretability: Circuits, Induction Headsmbrenndoerfer.com
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- What Can Transformers Learn In-Context?arxiv.org