flâneur

Scaling Activation Oracles to Trillion-Parameter Models

transluce.org · 7,403 words · saved by 3 readers

Many important AI behaviors, such as reward hacking or evaluation awareness, are difficult to detect purely from looking at an AI model's outputs. To help with this, we train activation oracles: AI assistants that analyze a model's internal activations to detect and predict that model's behavior. We scale the training of oracles to trillion-parameter models, and show that performance improves with model size, data size, and data quality. Our oracles achieve success on a range of difficult tasks such as predicting language switching and detecting reward hacking in coding agents.

This work is part of our ongoing efforts to train oversight foundation models: capable AI assistants that can help us to understand other AI models. We post updates like this one on the Oversight Foundations Blog. Many important AI behaviors, such as reward hacking or evaluation awareness, are difficult to detect purely from looking at an AI model's outputs. To help with this, we train activation oracles: AI assistants that analyze a model's internal activations to detect and predict that model's behavior. We scale the training of oracles to trillion-parameter models, and show that…

saved by

related reading