flâneur — a map of the web's best reading

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

arxiv.org · 14,108 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation Bowen Baker Joost Huizinga ∗† Leo Gao † Zehao Dou † Melody Y. Guan † Aleksander Madry † Wojciech Zaremba † Jakub Pachocki † David Farhi ∗† Core research team. Email correspondence to bowen@openai.com OpenAI Abstract Mitigating reward hacking—where AI systems misbehave due to flaws or misspecifications in their learning objectives—remains a key challenge in constructing capable and aligned models. We show that we can monitor a frontier reasoning model, such as OpenAI o3-mini, for reward hacking in agentic coding

Explore this link on the map →

saved by

related reading