CoT_Monitoring.pdf
cdn.openai.com · 8,949 words · saved by 1 readers
N/A
Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation Bowen Baker∗† Joost Huizinga∗† Leo Gao† Zehao Dou† Melody Y. Guan† Aleksander Madry† Wojciech Zaremba† Jakub Pachocki† David Farhi∗† Abstract Mitigating reward hacking—where AI systems misbehave due to flaws or misspecifications in their learning objectives—remains a key…
related reading
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- The Most Forbidden Technique — LessWronglesswrong.com
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- [2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv.org
- Training a Misaligned Reward Seekeralignment.anthropic.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com