flâneur

Sidestepping Evaluation Awareness and Anticipating Misalignment with Production Evaluations

alignment.openai.com · 3,157 words · saved by 1 readers

A pipeline to uncover unknown misaligned behavior and scale the creation of realistic evaluations.

← Back to OpenAI Alignment Blog Dec 18, 2025 · Marcus Williams, Cameron Raymond and Micah Carroll, in collaboration with the Safety Oversight team Evaluations of undesirable model behaviors are a critical tool for grounding current safety arguments. But evaluations best support claims about model safety when they cover the relevant risks and faithfully reflect real-world conditions. Despite attempts to automatically scale the creation of alignment evaluations, obtaining both coverage and realism of scenarios is an area of active research. In this post, we showcase a simple, scalable, and…

saved by

related reading