flâneur

CoT_Monitoring.pdf

cdn.openai.com · 8,949 words · saved by 1 readers

N/A

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation Bowen Baker∗† Joost Huizinga∗† Leo Gao† Zehao Dou† Melody Y. Guan† Aleksander Madry† Wojciech Zaremba† Jakub Pachocki† David Farhi∗† Abstract Mitigating reward hacking—where AI systems misbehave due to flaws or misspecifications in their learning objectives—remains a key…

related reading