✳flâneur — a map of the web's best reading
Metagaming matters for training, evaluation, and oversight
alignment.openai.com · 4,077 words · saved by 2 readers
Metagaming can complicate how we interpret behavior, and current models still give us a chance to study it directly.
Metagaming matters for training, evaluation, and oversight ← Back to OpenAI Alignment Blog Metagaming matters for training, evaluation, and oversight Mar 16, 2026 · Bronson Schoen (Apollo Research) and Jenny Nitishinskaya Metagaming can complicate how we interpret behavior, and current models still give us a chance to study it directly. As models become more capable, they also appear to become more situationally aware [ Laine , Needham ]. Some forms of situational awareness are undesirable and create risks. In current models, awareness of being in an evaluation has already influenc
Explore this link on the map →saved by
related reading
- A Toy Environment For Exploring Reasoning About Reward — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Models May Behave Worse When Eval Aware — LessWronglesswrong.com
- Predicting LLM Safety Before Release by Simulating Deploymentcdn.openai.com
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com
- How far does alignment midtraining generalize?alignment.openai.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Frontier Risk Report (February to March 2026) - METRmetr.org