How confessions can keep language models honest | OpenAI
We’re sharing an early, proof-of-concept method that trains models to report when they break instructions or take unintended shortcuts. AI systems are becoming more capable, and we want to understand them as deeply as possible—including how and why they arrive at an answer. Sometimes a model takes a shortcut or optimizes for the wrong objective, but its final output still looks correct. If we can surface when that happens, we can better monitor deployed systems, improve training, and increase trust in the outputs. Research by OpenAI and others has shown that AI models can hallucinate, reward-hack, or be dishonest. At the moment, we see the most concerning misbehaviors, such as scheming (opens in a new window) , only in stress-tests and adversarial evaluations. But as models become more capable and increasingly agentic, even rare forms of misalignment become more consequential, motivating us to invest in methods that help us better detect, understand, and mitigate these risks. This wo
December 3, 2025 Research Publication How confessions can keep language models honest We’re sharing an early, proof-of-concept method that trains models to report when they break instructions or take unintended shortcuts. Read the paper (opens in a new window) Loading… Share AI systems are becoming more capable, and we want to understand them as deeply as possible—including how and why they arrive at an answer. Sometimes a model takes a shortcut or optimizes for the wrong objective, but its final output still looks correct. If we can surface when that happens, we can better monitor deployed sy
Explore this link on the map →saved by
related reading
- confessions_paper.pdfcdn.openai.com
- Why we are excited about confession! — LessWronglesswrong.com
- Why We Are Excited About Confessionsalignment.openai.com
- Alignment faking in large language modelsarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- Models don’t seem to be dishonest in the way humans are — LessWronglesswrong.com
- How well do models follow their constitutions? — LessWronglesswrong.com
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com