Metagaming matters for training, evaluation, and oversight
alignment.openai.com · 4,077 words · saved by 2 readers
Metagaming can complicate how we interpret behavior, and current models still give us a chance to study it directly.
Metagaming matters for training, evaluation, and oversight ← Back to OpenAI Alignment Blog Metagaming matters for training, evaluation, and oversight Mar 16, 2026 · Bronson Schoen (Apollo Research) and Jenny Nitishinskaya Metagaming can complicate how we interpret behavior, and current models still give us a chance to study it directly. As models become more capable, they also appear to become more situationally aware [ Laine , Needham ]. Some forms of situational awareness are undesirable and create risks. In current models, awareness of being in an evaluation has already influenc
saved by
related reading
- A Toy Environment For Exploring Reasoning About Reward — LessWronglesswrong.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Why do models task game?greaterwrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- How far does alignment midtraining generalize?alignment.openai.com
- Frontier Risk Report (February to March 2026) - METRmetr.org
- Models May Behave Worse When Eval Aware — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- We need 3rd party Training-Run Assessments — LessWronglesswrong.com
- Predicting LLM Safety Before Release by Simulating Deploymentcdn.openai.com