Why we are excited about confession! — LessWrong
lesswrong.com · 8,476 words · saved by 1 readers
Boaz Barak, Gabriel Wu, Jeremy Chen, Manas Joglekar …
x Why we are excited about confession! — LessWrong AI Curated 2026 Top Fifty: 11 % 140 Why we are excited about confession! by Boaz Barak , Gabriel Wu , Manas Joglekar 14th Jan 2026 AI Alignment Forum Linkpost for alignment.openai.com 10 min read 32 140 Ω 54 Boaz Barak, Gabriel Wu, Jeremy Chen, Manas Joglekar [Linkposting from the OpenAI alignment blog , where we post more speculative/technical/informal results and thoughts on safety and alignment.] TL;DR We go into more details and some follow up results from our paper on confessions (see the original blog post ). We give deeper analysis of t
related reading
- Why We Are Excited About Confessionsalignment.openai.com
- How confessions can keep language models honest | OpenAIopenai.com
- confessions_paper.pdfcdn.openai.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Deep Deceptiveness — LessWronglesswrong.com
- The Most Forbidden Technique — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- Models don’t seem to be dishonest in the way humans are — LessWronglesswrong.com