Chapter 1: Transformer Interpretability - ARENA
Please send any problems / bugs on the #errata channel in the Slack group, and ask any questions on the dedicated channels for this chapter of material. If you want to change to dark mode, you can do this by clicking the three horizontal lines in the top-right, then navigating to Settings → Theme. Links to all other chapters: (0) Fundamentals, (1) Transformer Interpretability, (2) RL. This exercise set is built around linear probing, one of the most important tools in mechanistic interpretability for understanding what information language models represent internally. We'll look at three papers: The core idea: extract internal activations from a model, then train a simple classifier on them. If a linear probe can accurately classify some property from the activations, that property is linearly represented in the model's internal state. From the Geometry of Truth paper: "We identify a linear representation of truth that generalizes across several structurally and topically diverse datas
[1.3.1] Linear Probes Colab: exercises | solutions Please send any problems / bugs on the #errata channel in the Slack group , and ask any questions on the dedicated channels for this chapter of material. If you want to change to dark mode, you can do this by clicking the three horizontal lines in the top-right, then navigating to Settings → Theme. Links to all other chapters: (0) Fundamentals , (1) Transformer Interpretability , (2) RL . Introduction This exercise set is built around linear probing , one of the most important tools in mechanistic interpretability for understanding what inform
Explore this link on the map →saved by
related reading
- How well do truth probes generalise? — LessWronglesswrong.com
- Detecting Strategic Deception Using Linear Probes — LessWronglesswrong.com
- Transformer Circuits Threadtransformer-circuits.pub
- [2310.06824] The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasetsarxiv.org
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Simple probes can catch sleeper agents \ Anthropicanthropic.com
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Structure and Interpretation of Deep Networkssidn.baulab.info
- [2506.10805] Detecting High-Stakes Interactions with Activation Probesarxiv.org
- [2506.10805] Detecting High-Stakes Interactions with Activation Probesarxiv.org