✳flâneur — a map of the web's best reading
Why We Are Excited About Confessions
alignment.openai.com · 2,516 words · saved by 1 readers
A deeper look at confessions, reward hacking, and monitoring in alignment research.
Why We Are Excited About Confessions ← Back to OpenAI Alignment Blog Why We Are Excited About Confessions Jan 12, 2026 · Boaz Barak, Gabriel Wu, Jeremy Chen and Manas Joglekar TL;DR We go into more details and some follow up results from our paper on confessions (see the original blog post ). We give deeper analysis of the impact of training, as well as some preliminary comparisons to chain of thought monitoring. We have recently published a new paper on confessions, along with an accompanying blog post . Here, we want to share with the research community some of the reasons why we are excited
Explore this link on the map →related reading
- How confessions can keep language models honest | OpenAIopenai.com
- confessions_paper.pdfcdn.openai.com
- Why we are excited about confession! — LessWronglesswrong.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Alignment faking in large language modelsarxiv.org
- The Most Forbidden Technique — LessWronglesswrong.com
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com