Mechanistic Interpretability: Circuits, Induction Heads - Interactive | Michael Brenndoerfer
mbrenndoerfer.com · 6,280 words · saved by 1 readers
Reverse-engineer transformer networks into human-understandable algorithms by identifying circuits, induction heads, and mechanistic discoveries.
Mechanistic InterpretabilityLink Copied When a language model predicts that "The Eiffel Tower is located in" should be followed by "Paris," what computation produced that answer? Which weights fired, which attention heads activated, and which neurons collectively encoded the relevant geographic knowledge? Mechanistic interpretability is the research program that asks precisely these questions. Rather than treating a neural network as a black box that produces outputs from inputs, mechanistic interpretability tries to reverse-engineer the model into a human-understandable algorithm. It seeks…
saved by
related reading
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io
- Transformer Circuits Threadtransformer-circuits.pub
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Towards Automated Circuit Discovery for Mechanistic Interpretabilityarxiv.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- In-context Learning and Induction Headstransformer-circuits.pub
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumneelnanda.io
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Induction heads - illustrated — LessWronglesswrong.com