Gears-Level Mental Models of Transformer Interpretability — LessWrong
lesswrong.com · 2,207 words · saved by 1 readers
This post aims to quickly break down and explain the dominant mental models interpretability researchers currently use when thinking about how transf…
x Gears-Level Mental Models of Transformer Interpretability — LessWrong Interpretability (ML & AI) Language Models (LLMs) AI Frontpage 77 Gears-Level Mental Models of Transformer Interpretability by RowanWang 29th Mar 2022 AI Alignment Forum 7 min read 4 77 Ω 25 This post aims to quickly break down and explain the dominant mental models interpretability researchers currently use when thinking about how transformers work. In my view, the focus of transformer interpretability research is teleological: we care about the functions each component in the transformer performs and how those functions
related reading
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- Thinking like Transformersrush.github.io
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- An Analogy for Understanding Transformers — LessWronglesswrong.com
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- A Conceptual Guide to Transformers: Part Ibenlevinstein.substack.com
- Transformers from Scratche2eml.school
- Mechanistic Interpretability: Circuits, Induction Headsmbrenndoerfer.com
- In-context Learning and Induction Headstransformer-circuits.pub
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education