[2505.05410] Reasoning Models Don't Always Say What They Think
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2505.05410] Reasoning Models Don't Always Say What They Think Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Computation and Language arXiv:2505.05410 (cs) [Submitted on 8 May 2025] Title: Reasoning Models Don't Always Say What They Think Authors: Yanda Chen , Joe Benton , Ansh Radhakrishnan , Jonathan Uesato , Carson Denison , John Schulman , Arushi Somani , Peter Hase , Misha Wagner , Fabien Roger , Vlad Mikulik , Samuel R. Bowman , Jan Leike , Jared Kaplan , Ethan Perez View a
related reading
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- [2603.05706] Reasoning Models Struggle to Control their Chains of Thoughtarxiv.org
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- [2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv.org
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- [2510.27338] Reasoning Models Sometimes Output Illegible Chains of Thoughtarxiv.org
- What’s your AI thinking? - AI Digesttheaidigest.org
- Anthropic on X: "New Anthropic research: Do reasoning models accurately verbalize their reasoning? Our new paper shows they don't. This casts doubt on whether monitoring chains-of-thought (CoT) will be enough to reliably catch safety issuesx.com