✳flâneur — a map of the web's best reading
Discovering unknown AI misalignments in real-world usage
alignment.openai.com · 2,420 words · saved by 1 readers
Reasoning models can find and understand unknown misaligned behaviors from how users respond.
Discovering unknown AI misalignments in real-world usage ← Back to OpenAI Alignment Blog Discovering unknown AI misalignments in real-world usage Jan 2026 · Hannah Sheahan Reasoning models can find and understand unknown misaligned behaviors from how users respond. As AI systems become more capable and reach a broader audience, the range of real-world interactions—and the failures that accompany them—will continue to expand. Despite extensive testing, there is a limit to what we can learn in a lab before deployment. Inevitably, “unknown unknowns” arise only once models are exposed to the full
Explore this link on the map →related reading
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- How confessions can keep language models honest | OpenAIopenai.com
- Teaching Claude why \ Anthropicanthropic.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Alignment will happen by default. What’s next? — LessWronglesswrong.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Alignment faking in large language modelsarxiv.org
- How well do models follow their constitutions? — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com