flâneur — a map of the web's best reading

Discovering unknown AI misalignments in real-world usage

alignment.openai.com · 2,420 words · saved by 1 readers

Reasoning models can find and understand unknown misaligned behaviors from how users respond.

Discovering unknown AI misalignments in real-world usage ← Back to OpenAI Alignment Blog Discovering unknown AI misalignments in real-world usage Jan 2026 · Hannah Sheahan Reasoning models can find and understand unknown misaligned behaviors from how users respond. As AI systems become more capable and reach a broader audience, the range of real-world interactions—and the failures that accompany them—will continue to expand. Despite extensive testing, there is a limit to what we can learn in a lab before deployment. Inevitably, “unknown unknowns” arise only once models are exposed to the full

Explore this link on the map →

related reading