[2606.02609] Building Better Activation Oracles
Abstract:Activation Oracles (AOs) are promising methods for interpreting residual stream activations. However, current AOs face important issues, such as hallucinations and vagueness. Additionally, text-inversion confounds make them hard to evaluate. To this end, we improve the Activation Oracle (AO) training regime in four ways: training on on-policy rollouts, improving the conversational dataset, feeding more layers and an improvement to the injection formula. The capability improvements are marginal, but quality of life improvements are quite substantial. In addition, we open source the first comprehensive evaluation suite for AO quality, which we call AObench. Overall, we hope that our work sets a foundation that helps improve AOs and other models in the paradigm of scalable, end-to-end interpretability.
[2606.02609] Building Better Activation Oracles Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Machine Learning arXiv:2606.02609 (cs) [Submitted on 23 May 2026 ( v1 ), last revised 4 Jun 2026 (this version, v2)] Title: Building Better Activation Oracles Authors: Jan Bauer , Celeste De Schamphelaere , Adam Karvonen , Niclas Luick , Neel Nanda View a PDF of the paper titled Building Better Activation Oracles, by Jan Bauer and 4 other authors View PDF HTML (experimental) Abstract: Ac
Explore this link on the map →related reading
- [2606.02609] Building Better Activation Oraclesarxiv.org
- Current activation oracles are hard to use — LessWronglesswrong.com
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers — LessWronglesswrong.com
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainersalignment.anthropic.com
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- cot-oracle/AObench at main · ceselder/cot-oracle · GitHubgithub.com
- GitHub - ceselder/cot-oracle: Our improvements to activation oracles · GitHubgithub.com
- [2512.15674] Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainersarxiv.org
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Dreaming Vectors: Gradient-descented steering vectors from Activation Oracles and using them to Red-Team AOs — LessWronglesswrong.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Natural Language Autoencoders \ Anthropicanthropic.com