Scaling Activation Oracles to Trillion-Parameter Models
Many important AI behaviors, such as reward hacking or evaluation awareness, are difficult to detect purely from looking at an AI model's outputs. To help with this, we train activation oracles: AI assistants that analyze a model's internal activations to detect and predict that model's behavior. We scale the training of oracles to trillion-parameter models, and show that performance improves with model size, data size, and data quality. Our oracles achieve success on a range of difficult tasks such as predicting language switching and detecting reward hacking in coding agents.
This work is part of our ongoing efforts to train oversight foundation models: capable AI assistants that can help us to understand other AI models. We post updates like this one on the Oversight Foundations Blog. Many important AI behaviors, such as reward hacking or evaluation awareness, are difficult to detect purely from looking at an AI model's outputs. To help with this, we train activation oracles: AI assistants that analyze a model's internal activations to detect and predict that model's behavior. We scale the training of oracles to trillion-parameter models, and show that…
saved by
related reading
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- [2606.02609] Building Better Activation Oraclesarxiv.org
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- Foundation Models for Oversight | Transluce AItransluce.org
- Current activation oracles are hard to use — LessWronglesswrong.com
- [2606.02609] Building Better Activation Oraclesarxiv.org
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainersalignment.anthropic.com
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- [2512.15674] Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainersarxiv.org