In-context Learning and Induction Heads
As Transformer generative models continue to scale and gain increasing real world use , addressing their associated safety problems becomes increasingly important. Mechanistic interpretability – attempting to reverse engineer the detailed computations performed by the model – offers one possible avenue for addressing these safety issues. If we can understand the internal structures that cause Transformer models to produce the outputs they do, then we may be able to address current safety problems more systematically, as well as anticipating safety problems in future more powerful models. In the past, mechanistic interpretability has largely focused on CNN vision models, but recently, we presented some very preliminary progress on mechanistic interpretability for Transformer language models. Specifically, in our prior work we developed a mathematical framework for decomposing the operations of transformers, which allowed us to make sense of small (1 and 2 layer attention-only) models
In-context Learning and Induction Heads Transformer Circuits Thread In-context Learning and Induction Heads Authors Catherine Olsson ∗ , Nelson Elhage ∗ , Neel Nanda ∗ , Nicholas Joseph † , Nova DasSarma † , Tom Henighan † , Ben Mann † , Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds , Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, Chris Olah ‡ Affiliation Anthropic Published Mar 8, 2022 * Core Research Contributor; † Core Infrastructur
Explore this link on the map →saved by
- Winnie Xu
- Amir
- Varun Shenoy
- Claire Wang
- Asma Lamgh
- Emma Guo
- Derek Yen
- Lydia Nottingham
- Timothy Kostolansky
- Julian H
- Akira Yoshiyama
related reading
- Transformer Circuits Threadtransformer-circuits.pub
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Induction heads - illustrated — LessWronglesswrong.com
- [2212.07677] Transformers learn in-context by gradient descentarxiv.org
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- Softmax Linear Unitstransformer-circuits.pub
- Transformer (deep learning) - Wikipediaen.wikipedia.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- How does in-context learning work? A framework for understanding the differences from traditional supervised learning | SAIL Blogai.stanford.edu