Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers — LessWrong
TL;DR: We train LLMs to accept LLM neural activations as inputs and answer arbitrary questions about them in natural language. These Activation Oracl…
x Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers — LessWrong Interpretability (ML & AI) Language Models (LLMs) AI Frontpage 2025 Top Fifty: 14 % 154 Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers by Sam Marks , Adam Karvonen , James Chua , Subhash Kantamneni , Euan Ong , Julian Minder , Clément Dumas , Owain_Evans 18th Dec 2025 AI Alignment Forum Linkpost for arxiv.org 10 min read 11 154 Ω 73 TL;DR: We train LLMs to accept LLM neural activations as inputs and answer arbitrary questions about them in natural l
Explore this link on the map →related reading
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainersalignment.anthropic.com
- [2512.15674] Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainersarxiv.org
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Current activation oracles are hard to use — LessWronglesswrong.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Transformer Circuits Threadtransformer-circuits.pub
- [2606.02609] Building Better Activation Oraclesarxiv.org
- [2606.02609] Building Better Activation Oraclesarxiv.org
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWronglesswrong.com
- Dreaming Vectors: Gradient-descented steering vectors from Activation Oracles and using them to Red-Team AOs — LessWronglesswrong.com