flâneur — a map of the web's best reading

Metagaming matters for training, evaluation, and oversight

alignment.openai.com · 4,077 words · saved by 2 readers

Metagaming can complicate how we interpret behavior, and current models still give us a chance to study it directly.

Metagaming matters for training, evaluation, and oversight ← Back to OpenAI Alignment Blog Metagaming matters for training, evaluation, and oversight Mar 16, 2026 · Bronson Schoen (Apollo Research) and Jenny Nitishinskaya Metagaming can complicate how we interpret behavior, and current models still give us a chance to study it directly. As models become more capable, they also appear to become more situationally aware [ Laine , Needham ]. Some forms of situational awareness are undesirable and create risks. In current models, awareness of being in an evaluation has already influenc

Explore this link on the map →

saved by

related reading