✳flâneur — a map of the web's best reading
Why we are excited about confession! — LessWrong
lesswrong.com · 8,476 words · saved by 1 readers
Boaz Barak, Gabriel Wu, Jeremy Chen, Manas Joglekar …
x Why we are excited about confession! — LessWrong AI Curated 2026 Top Fifty: 11 % 140 Why we are excited about confession! by Boaz Barak , Gabriel Wu , Manas Joglekar 14th Jan 2026 AI Alignment Forum Linkpost for alignment.openai.com 10 min read 32 140 Ω 54 Boaz Barak, Gabriel Wu, Jeremy Chen, Manas Joglekar [Linkposting from the OpenAI alignment blog , where we post more speculative/technical/informal results and thoughts on safety and alignment.] TL;DR We go into more details and some follow up results from our paper on confessions (see the original blog post ). We give deeper analysis of t
Explore this link on the map →related reading
- How confessions can keep language models honest | OpenAIopenai.com
- confessions_paper.pdfcdn.openai.com
- Why We Are Excited About Confessionsalignment.openai.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Alignment faking in large language modelsarxiv.org
- The Most Forbidden Technique — LessWronglesswrong.com