Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation Bowen Baker Joost Huizinga ∗† Leo Gao † Zehao Dou † Melody Y. Guan † Aleksander Madry † Wojciech Zaremba † Jakub Pachocki † David Farhi ∗† Core research team. Email correspondence to bowen@openai.com OpenAI Abstract Mitigating reward hacking—where AI systems misbehave due to flaws or misspecifications in their learning objectives—remains a key challenge in constructing capable and aligned models. We show that we can monitor a frontier reasoning model, such as OpenAI o3-mini, for reward hacking in agentic coding
Explore this link on the map →saved by
related reading
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- The Most Forbidden Technique — LessWronglesswrong.com
- [Research Note] Optimizing The Final Output Can Obfuscate CoT — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- [2512.18311] Monitoring Monitorabilityarxiv.org
- [2603.05706] Reasoning Models Struggle to Control their Chains of Thoughtarxiv.org
- [2512.18311] Monitoring Monitorabilityarxiv.org
- [2505.05410] Reasoning Models Don't Always Say What They Thinkarxiv.org
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com