Why We Are Excited About Confessions
alignment.openai.com · 2,516 words · saved by 3 readers
A deeper look at confessions, reward hacking, and monitoring in alignment research.
Why We Are Excited About Confessions ← Back to OpenAI Alignment Blog Why We Are Excited About Confessions Jan 12, 2026 · Boaz Barak, Gabriel Wu, Jeremy Chen and Manas Joglekar TL;DR We go into more details and some follow up results from our paper on confessions (see the original blog post ). We give deeper analysis of the impact of training, as well as some preliminary comparisons to chain of thought monitoring. We have recently published a new paper on confessions, along with an accompanying blog post . Here, we want to share with the research community some of the reasons why we are excited
saved by
related reading
- How confessions can keep language models honest | OpenAIopenai.com
- confessions_paper.pdfcdn.openai.com
- Why we are excited about confession! — LessWronglesswrong.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- [2510.27062] Consistency Training Helps Stop Sycophancy and Jailbreaksarxiv.org
- Self-CTRL: Self-Consistency Training with Reinforcement Learningarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- How AI Is Learning to Think in Secretnickandresen.substack.com
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org