Chapter 1: Transformer Interpretability - ARENA
Please send any problems / bugs on the #errata channel in the Slack group, and ask any questions on the dedicated channels for this chapter of material. If you want to change to dark mode, you can do this by clicking the three horizontal lines in the top-right, then navigating to Settings → Theme. Links to all other chapters: (0) Fundamentals, (1) Transformer Interpretability, (2) RL. These pages are designed to get you introduced to the core concepts of mechanistic interpretability, via Neel Nanda's TransformerLens library. Most of the sections are constructed in the following way: The running theme of the exercises is induction circuits. Induction circuits are a particular type of circuit in a transformer, which can perform basic in-context learning. You should read the corresponding section of Neel's glossary, before continuing. This LessWrong post might also help; it contains some diagrams (like the one below) which walk through the induction mechanism step by step. Each exercise w
[1.2] Intro to Mechanistic Interpretability: TransformerLens & induction circuits Colab: exercises | solutions Please send any problems / bugs on the #errata channel in the Slack group , and ask any questions on the dedicated channels for this chapter of material. If you want to change to dark mode, you can do this by clicking the three horizontal lines in the top-right, then navigating to Settings → Theme. Links to all other chapters: (0) Fundamentals , (1) Transformer Interpretability , (2) RL . Introduction These pages are designed to get you introduced to the core concepts of mechanistic i
Explore this link on the map →related reading
- Transformer Circuits Threadtransformer-circuits.pub
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- In-context Learning and Induction Headstransformer-circuits.pub
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Exploratory Analysis Demo - TransformerLens Documentationtransformerlensorg.github.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- TransformerLens Documentationtransformerlensorg.github.io
- TransformerLens Documentationtransformerlensorg.github.io
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io