When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
\uselogo \correspondingauthor semmons@google.com \reportnumber 186324 When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors Scott Emmons \thepa Erik Jenner \thepa David K. Elson \thepa Rif A. Saurous Google, Paradigms of Intelligence Team Senthooran Rajamanoharan \thepa Heng Chen \thepa Irhum Shafkat \thepa Rohin Shah \thepa Abstract While chain-of-thought (CoT) monitoring is an appealing AI safety defense, recent work on “unfaithfulness” has cast doubt on its reliability. These findings highlight an important failure mode, particularly when CoT acts as a post-hoc rati
Explore this link on the map →saved by
related reading
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- What’s your AI thinking? - AI Digesttheaidigest.org
- [2603.05706] Reasoning Models Struggle to Control their Chains of Thoughtarxiv.org
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- [2505.05410] Reasoning Models Don't Always Say What They Thinkarxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- How AI Is Learning to Think in Secret — LessWronglesswrong.com
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- [2512.18311] Monitoring Monitorabilityarxiv.org