✳flâneur — a map of the web's best reading
Gears-Level Mental Models of Transformer Interpretability — LessWrong
lesswrong.com · 2,207 words · saved by 1 readers
This post aims to quickly break down and explain the dominant mental models interpretability researchers currently use when thinking about how transf…
x Gears-Level Mental Models of Transformer Interpretability — LessWrong Interpretability (ML & AI) Language Models (LLMs) AI Frontpage 77 Gears-Level Mental Models of Transformer Interpretability by RowanWang 29th Mar 2022 AI Alignment Forum 7 min read 4 77 Ω 25 This post aims to quickly break down and explain the dominant mental models interpretability researchers currently use when thinking about how transformers work. In my view, the focus of transformer interpretability research is teleological: we care about the functions each component in the transformer performs and how those functions
Explore this link on the map →related reading
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- An Analogy for Understanding Transformers — LessWronglesswrong.com
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Transformers from Scratche2eml.school
- A Conceptual Guide to Transformers: Part Ibenlevinstein.substack.com
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- In-context Learning and Induction Headstransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org