[2603.05706] Reasoning Models Struggle to Control their Chains of Thought
Abstract:Chain-of-thought (CoT) monitoring is a promising tool for detecting misbehaviors and understanding the motivations of modern reasoning models. However, if models can control what they verbalize in their CoT, it could undermine CoT monitorability. To measure this undesirable capability -- CoT controllability -- we introduce the CoT-Control evaluation suite, which includes tasks that require models to solve problems while adhering to CoT instructions, e.g., reasoning about a genetics question without using the word 'chromosome'. We show that reasoning models possess significantly lower CoT controllability than output controllability; for instance, Claude Sonnet 4.5 can control its CoT only 2.7% of the time but 61.9% when controlling its final output. We also find that CoT controllability is higher for larger models and decreases with more RL training, test-time compute, and increased problem difficulty. CoT controllability failures extend even to situations in which models are given incentives (as opposed to direct requests) to evade CoT monitors, although models exhibit slightly higher controllability when they are told they are being monitored. Similarly, eliciting controllability by adversarially optimizing prompts does not meaningfully increase controllability. Our results leave us cautiously optimistic that CoT controllability is currently unlikely to be a failure mode of CoT monitorability. However, the mechanism behind low controllability is not well understood. Given its importance for maintaining CoT monitorability, we recommend that frontier labs track CoT controllability in future models.
Reasoning Models Struggle to Control their Chains of Thought Chen Yueh-Han∗ Robert McCarthy Bruce W. Lee He He NYU, MATS UCL, MATS UPenn, MATS NYU Ian Kivlichan Bowen Baker Micah Carroll Tomek Korbak∗ OpenAI…
saved by
related reading
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- [2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv.org
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- [2505.05410] Reasoning Models Don't Always Say What They Thinkarxiv.org
- Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settingsarxiv.org
- What’s your AI thinking? - AI Digesttheaidigest.org
- Policy Options for Preserving Chain of Thought Monitorability — Institute for AI Policy and Strategyiaps.ai
- [2512.18311] Monitoring Monitorabilityarxiv.org
- [2510.27338] Reasoning Models Sometimes Output Illegible Chains of Thoughtarxiv.org
- [2512.18311] Monitoring Monitorabilityarxiv.org