Sidestepping Evaluation Awareness and Anticipating Misalignment with Production Evaluations
alignment.openai.com · 3,157 words · saved by 1 readers
A pipeline to uncover unknown misaligned behavior and scale the creation of realistic evaluations.
← Back to OpenAI Alignment Blog Dec 18, 2025 · Marcus Williams, Cameron Raymond and Micah Carroll, in collaboration with the Safety Oversight team Evaluations of undesirable model behaviors are a critical tool for grounding current safety arguments. But evaluations best support claims about model safety when they cover the relevant risks and faithfully reflect real-world conditions. Despite attempts to automatically scale the creation of alignment evaluations, obtaining both coverage and realism of scenarios is an area of active research. In this post, we showcase a simple, scalable, and…
saved by
related reading
- Toward A Public Science of Model Behavior | Transluce AItransluce.org
- We need 3rd party Training-Run Assessments — LessWronglesswrong.com
- Predicting LLM Safety Before Release by Simulating Deploymentcdn.openai.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Realistic Evaluations Will Not Prevent Evaluation Awareness — LessWronglesswrong.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Predicting model behavior before release by simulating deployment | OpenAIopenai.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Can public chat data predict real-world AI misalignments?alignment.openai.com
- Teaching Claude Whyalignment.anthropic.com
- How far does alignment midtraining generalize?alignment.openai.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com