Anthropic on X: "New Anthropic research: Do reasoning models accurately verbalize their reasoning? Our new paper shows they don't. This casts doubt on whether monitoring chains-of-thought (CoT) will be enough to reliably catch safety issues. https://t.co/K3MrwqUXX9" / X
To view keyboard shortcuts, press question mark View keyboard shortcuts Messages Home Explore Notifications Messages Grok Premium Lists Bookmarks Jobs Communities Verified Orgs Profile More Post Tasha @TashaPais Post Reply See new posts Conversation Chris Olah reposted Anthropic @AnthropicAI New Anthropic research: Do reasoning models accurately verbalize their reasoning? Our new paper shows they don't. This casts doubt on whether monitoring chains-of-thought (CoT) will be enough to reliably catch safety issues. ALT 9:31 AM · Apr 3, 2025 · 1M Views 146 839 3.6K 1.5K Post your reply Reply Anthropic @AnthropicAI · Apr 3 We slipped problem-solving hints to Claude 3.7 Sonnet and DeepSeek R1, then tested whether their Chains-of-Thought would mention using the hint (if the models actually used it). Read the blog: https://anthropic.com/research/reasoning-models-dont-say-think… 7 11 197 16K Anthropic @AnthropicAI · Apr 3 We found Chains-of-Thought largely aren’t “faithful”: the rate of m
Anthropic @AnthropicAI New Anthropic research: Do reasoning models accurately verbalize their reasoning? Our new paper shows they don't. This casts doubt on whether monitoring chains-of-thought (CoT) will be enough to reliably catch safety issues. 4:31 PM · Apr 3, 2025 1.1M Views 146 0 1 4 6 566 0 5 6 6 3.5K 0 3 . 5 K 1.5K 0 1 . 5 K Read 146 replies
Explore this link on the map →related reading
- [2505.05410] Reasoning Models Don't Always Say What They Thinkarxiv.org
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- [2603.05706] Reasoning Models Struggle to Control their Chains of Thoughtarxiv.org
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- Measuring Faithfulness in Chain-of-Thought Reasoning \ Anthropicanthropic.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- [2510.27338] Reasoning Models Sometimes Output Illegible Chains of Thoughtarxiv.org
- What’s your AI thinking? - AI Digesttheaidigest.org
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org