Reasoning models don't always say what they think \ Anthropic
Since late last year, “reasoning models” have been everywhere. These are AI models—such as Claude 3.7 Sonnet—that show their working: as well as their eventual answer, you can read the (often fascinating and convoluted) way that they got there, in what’s called their “Chain-of-Thought”. As well as helping reasoning models work their way through more difficult problems, the Chain-of-Thought has been a boon for AI safety researchers. That’s because we can (among other things) check for things the model says in its Chain-of-Thought that go unsaid in its output, which can help us spot undesirable behaviours like deception. But if we want to use the Chain-of-Thought for alignment purposes, there’s a crucial question: can we actually trust what models say in their Chain-of-Thought? In a perfect world, everything in the Chain-of-Thought would be both understandable to the reader, and it would be faithful—it would be a true description of exactly what the model was thinking as it reached its a
Alignment Reasoning models don't always say what they think Apr 3, 2025 Read the paper Since late last year, “reasoning models” have been everywhere. These are AI models—such as Claude 3.7 Sonnet—that show their working : as well as their eventual answer, you can read the (often fascinating and convoluted) way that they got there, in what’s called their “Chain-of-Thought”. As well as helping reasoning models work their way through more difficult problems, the Chain-of-Thought has been a boon for AI safety researchers. That’s because we can (among other things) check for things the model says i
saved by
related reading
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- [2505.05410] Reasoning Models Don't Always Say What They Thinkarxiv.org
- Measuring Faithfulness in Chain-of-Thought Reasoning \ Anthropicanthropic.com
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- Is AI Reasoning Right for the Wrong Reasons? | Quanta Magazinequantamagazine.org
- What’s your AI thinking? - AI Digesttheaidigest.org
- How AI Is Learning to Think in Secretnickandresen.substack.com
- Why We Are Excited About Confessionsalignment.openai.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- How AI Is Learning to Think in Secret — LessWronglesswrong.com