When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
\uselogo \correspondingauthor semmons@google.com \reportnumber 186324 When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors Scott Emmons \thepa Erik Jenner \thepa David K. Elson \thepa Rif A. Saurous Google, Paradigms of Intelligence Team Senthooran Rajamanoharan \thepa Heng Chen \thepa Irhum Shafkat \thepa Rohin Shah \thepa Abstract While chain-of-thought (CoT) monitoring is an appealing AI safety defense, recent work on “unfaithfulness” has cast doubt on its reliability. These findings highlight an important failure mode, particularly when CoT acts as a post-hoc rati
saved by
related reading
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- [2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv.org
- [2603.05706] Reasoning Models Struggle to Control their Chains of Thoughtarxiv.org
- What’s your AI thinking? - AI Digesttheaidigest.org
- Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settingsarxiv.org
- [2505.05410] Reasoning Models Don't Always Say What They Thinkarxiv.org
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- [2512.11949] Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitorsarxiv.org
- Policy Options for Preserving Chain of Thought Monitorability — Institute for AI Policy and Strategyiaps.ai
- How AI Is Learning to Think in Secret — LessWronglesswrong.com