[2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
Abstract:AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known AI oversight methods, CoT monitoring is imperfect and allows some misbehavior to go unnoticed. Nevertheless, it shows promise and we recommend further research into CoT monitorability and investment in CoT monitoring alongside existing safety methods. Because CoT monitorability may be fragile, we recommend that frontier model developers consider the impact of development decisions on CoT monitorability.
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety Tomek Korbak∗ UK AI Security Institute Mikita Balesni∗ Apollo Research Elizabeth Barnes METR Yoshua Bengio University of Montreal & Mila Joe Benton Anthropic Joseph Bloom UK AI Security Institute arXiv:2507.11473v2 [cs.AI] 7…
saved by
related reading
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- Policy Options for Preserving Chain of Thought Monitorability — Institute for AI Policy and Strategyiaps.ai
- [2512.11949] Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitorsarxiv.org
- How AI Is Learning to Think in Secretnickandresen.substack.com
- SPAR Research Librarylibrary.sparai.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- What’s your AI thinking? - AI Digesttheaidigest.org
- [2603.05706] Reasoning Models Struggle to Control their Chains of Thoughtarxiv.org