[2505.05410] Reasoning Models Don't Always Say What They Think
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2505.05410] Reasoning Models Don't Always Say What They Think Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Computation and Language arXiv:2505.05410 (cs) [Submitted on 8 May 2025] Title: Reasoning Models Don't Always Say What They Think Authors: Yanda Chen , Joe Benton , Ansh Radhakrishnan , Jonathan Uesato , Carson Denison , John Schulman , Arushi Somani , Peter Hase , Misha Wagner , Fabien Roger , Vlad Mikulik , Samuel R. Bowman , Jan Leike , Jared Kaplan , Ethan Perez View a
Explore this link on the map →related reading
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- [2603.05706] Reasoning Models Struggle to Control their Chains of Thoughtarxiv.org
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- [2510.27338] Reasoning Models Sometimes Output Illegible Chains of Thoughtarxiv.org
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- What’s your AI thinking? - AI Digesttheaidigest.org
- Anthropic on X: "New Anthropic research: Do reasoning models accurately verbalize their reasoning? Our new paper shows they don't. This casts doubt on whether monitoring chains-of-thought (CoT) will be enough to reliably catch safety issuesx.com
- Measuring Faithfulness in Chain-of-Thought Reasoning \ Anthropicanthropic.com