Policy Options for Preserving Chain of Thought Monitorability — Institute for AI Policy and Strategy
By using this website, you agree to our use of cookies. We use cookies to provide you with a great experience and to help our website run effectively. This report was co-authored by Oscar Delaney, Oliver Guest, and Renan Araujo. Today’s most advanced artificial intelligence (AI) models use “chain of thought” (CoT) reasoning; monitoring this CoT can be valuable for controlling these systems and ensuring that they will behave as intended. This is because, like humans, AIs need to think in steps in order to solve complex problems. We can monitor the steps that the model writes (i.e., the CoT) and intervene to stop the model’s default course of action, if a concerning intention is expressed in the CoT. As AI systems are deployed in higher stakes contexts, CoT monitoring could become increasingly important for preventing risks to public safety and national security. In experiments, researchers have already found AIs to cheat on a software development task by altering a test to make the exis
This report was co-authored by Oscar Delaney, Oliver Guest, and Renan Araujo. Executive Summary Today’s most advanced artificial intelligence (AI) models use “chain of thought” (CoT) reasoning; monitoring this CoT can be valuable for controlling these systems and ensuring that they will behave as intended. This is because, like humans, AIs need to think in steps in order to solve complex problems. We can monitor the steps that the model writes (i.e., the CoT) and intervene to stop the model’s default course of action, if a concerning intention is expressed in the CoT. As AI systems are…
saved by
related reading
- [2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv.org
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- [2603.05706] Reasoning Models Struggle to Control their Chains of Thoughtarxiv.org
- What’s your AI thinking? - AI Digesttheaidigest.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- The Most Forbidden Technique — LessWronglesswrong.com
- [2512.18311] Monitoring Monitorabilityarxiv.org
- [2512.18311] Monitoring Monitorabilityarxiv.org