Chapter 1: Transformer Interpretability - ARENA
Please send any problems / bugs on the #errata channel in the Slack group, and ask any questions on the dedicated channels for this chapter of material. If you want to change to dark mode, you can do this by clicking the three horizontal lines in the top-right, then navigating to Settings → Theme. Links to all other chapters: (0) Fundamentals, (1) Transformer Interpretability, (2) RL. Linear probes let you ask yes-or-no questions about a model's activations, and SAEs give you an unsupervised decomposition into interpretable features. Both are useful, but they share a limitation: you have to decide what you're looking for before you start looking. What if you could just ask a model's activations an open-ended question in plain English? "What concept is this layer encoding?" or "Is this model planning to lie?" That's the idea behind Activation Oracles: LLMs that have been trained to take another model's internal activations as input and answer arbitrary questions about them. The oracle r
[1.3.4] Activation Oracles Colab: exercises | solutions Please send any problems / bugs on the #errata channel in the Slack group , and ask any questions on the dedicated channels for this chapter of material. If you want to change to dark mode, you can do this by clicking the three horizontal lines in the top-right, then navigating to Settings → Theme. Links to all other chapters: (0) Fundamentals , (1) Transformer Interpretability , (2) RL . Introduction Linear probes let you ask yes-or-no questions about a model's activations, and SAEs give you an unsupervised decomposition into interpretab
Explore this link on the map →related reading
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers — LessWronglesswrong.com
- Transformer Circuits Threadtransformer-circuits.pub
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainersalignment.anthropic.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- [2606.02609] Building Better Activation Oraclesarxiv.org
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Natural Language Autoencoders \ Anthropicanthropic.com
- Current activation oracles are hard to use — LessWronglesswrong.com
- [2606.02609] Building Better Activation Oraclesarxiv.org
- Dreaming Vectors: Gradient-descented steering vectors from Activation Oracles and using them to Red-Team AOs — LessWronglesswrong.com
- [2512.15674] Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainersarxiv.org