Chapter 1: Transformer Interpretability - ARENA
Please send any problems / bugs on the #errata channel in the Slack group, and ask any questions on the dedicated channels for this chapter of material. If you want to change to dark mode, you can do this by clicking the three horizontal lines in the top-right, then navigating to Settings → Theme. Links to all other chapters: (0) Fundamentals, (1) Transformer Interpretability, (2) RL. This notebook / document is built around the Interpretability in the Wild paper, in which the authors aim to understand the indirect object identification circuit in GPT-2 small. This circuit is responsible for the model's ability to complete sentences like "John and Mary went to the shops, John gave a bag to" with the correct token "" Mary". It is loosely divided into different sections, each one with their own flavour. Sections 1, 2 & 3 are derived from Neel Nanda's notebook Exploratory_Analysis_Demo. The flavour of these exercises is experimental and loose, with a focus on demonstrating what explorator
[1.4.1] Indirect Object Identification Colab: exercises | solutions Please send any problems / bugs on the #errata channel in the Slack group , and ask any questions on the dedicated channels for this chapter of material. If you want to change to dark mode, you can do this by clicking the three horizontal lines in the top-right, then navigating to Settings → Theme. Links to all other chapters: (0) Fundamentals , (1) Transformer Interpretability , (2) RL . Introduction This notebook / document is built around the Interpretability in the Wild paper, in which the authors aim to understand the ind
Explore this link on the map →related reading
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Exploratory Analysis Demo - TransformerLens Documentationtransformerlensorg.github.io
- Transformer Circuits Threadtransformer-circuits.pub
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Attribution Patching: Activation Patching At Industrial Scale - Neel Nandaneelnanda.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- In-context Learning and Induction Headstransformer-circuits.pub
- Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org