Chapter 1: Transformer Interpretability - ARENA
Please send any problems / bugs on the #errata channel in the Slack group, and ask any questions on the dedicated channels for this chapter of material. If you want to change to dark mode, you can do this by clicking the three horizontal lines in the top-right, then navigating to Settings → Theme. Links to all other chapters: (0) Fundamentals, (1) Transformer Interpretability, (2) RL. This exercise set is built around linear probing, one of the most important tools in mechanistic interpretability for understanding what information language models represent internally. We'll look at three papers: The core idea: extract internal activations from a model, then train a simple classifier on them. If a linear probe can accurately classify some property from the activations, that property is linearly represented in the model's internal state. From the Geometry of Truth paper: "We identify a linear representation of truth that generalizes across several structurally and topically diverse datas
[1.3.1] Linear Probes Colab: exercises | solutions Please send any problems / bugs on the #errata channel in the Slack group , and ask any questions on the dedicated channels for this chapter of material. If you want to change to dark mode, you can do this by clicking the three horizontal lines in the top-right, then navigating to Settings → Theme. Links to all other chapters: (0) Fundamentals , (1) Transformer Interpretability , (2) RL . Introduction This exercise set is built around linear probing , one of the most important tools in mechanistic interpretability for understanding what inform
saved by
related reading
- How well do truth probes generalise? — LessWronglesswrong.com
- Detecting Strategic Deception Using Linear Probes — LessWronglesswrong.com
- [2310.06824] The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasetsarxiv.org
- Transformer Circuits Threadtransformer-circuits.pub
- Structure and Interpretation of Deep Networkssidn.baulab.info
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- [2607.14111] Introspection Fine-Tuning (IFT): Training Small LLMs to Introspectarxiv.org
- Actually, Othello-GPT Has A Linear Emergent World Representation - Neel Nandaneelnanda.io
- The Linear Representation Hypothesis and the Geometry of Large Language Modelsarxiv.org
- Topicslearnmechinterp.com
- [2506.10805] Detecting High-Stakes Interactions with Activation Probesarxiv.org