Reasoning models don't always say what they think \ Anthropic
Since late last year, “reasoning models” have been everywhere. These are AI models—such as Claude 3.7 Sonnet—that show their working: as well as their eventual answer, you can read the (often fascinating and convoluted) way that they got there, in what’s called their “Chain-of-Thought”. As well as helping reasoning models work their way through more difficult problems, the Chain-of-Thought has been a boon for AI safety researchers. That’s because we can (among other things) check for things the model says in its Chain-of-Thought that go unsaid in its output, which can help us spot undesirable behaviours like deception. But if we want to use the Chain-of-Thought for alignment purposes, there’s a crucial question: can we actually trust what models say in their Chain-of-Thought? In a perfect world, everything in the Chain-of-Thought would be both understandable to the reader, and it would be faithful—it would be a true description of exactly what the model was thinking as it reached its a
Alignment Reasoning models don't always say what they think Apr 3, 2025 Read the paper Since late last year, “reasoning models” have been everywhere. These are AI models—such as Claude 3.7 Sonnet—that show their working : as well as their eventual answer, you can read the (often fascinating and convoluted) way that they got there, in what’s called their “Chain-of-Thought”. As well as helping reasoning models work their way through more difficult problems, the Chain-of-Thought has been a boon for AI safety researchers. That’s because we can (among other things) check for things the model says i
Explore this link on the map →saved by
related reading
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- [2505.05410] Reasoning Models Don't Always Say What They Thinkarxiv.org
- Measuring Faithfulness in Chain-of-Thought Reasoning \ Anthropicanthropic.com
- How confessions can keep language models honest | OpenAIopenai.com
- What’s your AI thinking? - AI Digesttheaidigest.org
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- How AI Is Learning to Think in Secret — LessWronglesswrong.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Towards Faithful Chain-of-Thought: Large Language Models are Bridging Reasonersarxiv.org
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org