✳flâneur — a map of the web's best reading
confessions_paper.pdf
cdn.openai.com · 21,955 words · saved by 2 readers
N/A
# link_1lavq6x4dpr.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=true - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - CreationDate=D:20251203012233Z - Creator=LaTeX with hyperref - ModDate=D:20251203012233Z - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.27 (TeX Live 2025) kpathsea version 6.4.1 - Producer=pdfTeX-1.40.27 - Trapped=False ## Contents ### Page 1 Training LLMs for Honesty via ConfessionsManas Joglekar∗ Jeremy Chen∗ Gabriel Wu∗Jason Yosinski Jasmine Wang Boaz Barak† Amelia Glaese†OpenAIAbstrac
Explore this link on the map →saved by
related reading
- How confessions can keep language models honest | OpenAIopenai.com
- Why We Are Excited About Confessionsalignment.openai.com
- Why we are excited about confession! — LessWronglesswrong.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Models don’t seem to be dishonest in the way humans are — LessWronglesswrong.com
- How well do models follow their constitutions? — LessWronglesswrong.com
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Sandbagging with misaligned action - Chain-of-Thought Transcript - Anti-Schemingantischeming.ai
- Agentic Misalignment in Summer 2026alignment.anthropic.com