What’s your AI thinking? - AI Digest
AI is rapidly becoming more capable – the time horizon for coding tasks is doubling every 4-7 months. But we don’t actually know what these increasingly capable models are thinking. And that’s a problem. If we can’t tell what a model is thinking, then we can’t tell when it is downplaying its capabilities, cheating on tests, or straight up working against us. Luckily we do have a lead: the chain of thought (CoT). This CoT is used in all top-performing language models. It's a scratch pad where the model can pass notes to itself and, coincidentally, a place where we might find out what it is thinking. Except, the CoT isn’t always faithful. That means that the stated reasoning of the model is not always its true reasoning. And we are not sure yet how to improve that. However, some researchers now argue that we don’t need complete faithfulness. They argue monitorability is sufficient. While faithfulness means you can read the model’s mind and know what it is thinking. Monitorability means y
What’s your AI thinking? - AI Digest Demos and explainers About Join the team What’s your AI thinking? What is Chain of Thought? Chain of Thought Faithfulness Chain of Thought Monitorability Keeping the Chain of Thought Monitorable What’s your AI thinking? A step by step introduction to chain of thought monitorability August 4, 2025 Shoshannah Tekofsky Member of technical staff at AI Digest AI is rapidly becoming more capable – the time horizon for coding tasks is doubling every 4-7 months . But we don’t actually know what these increasingly capable models are thinking . And that’s a problem.
related reading
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- [2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv.org
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- How AI Is Learning to Think in Secret — LessWronglesswrong.com
- How AI Is Learning to Think in Secretnickandresen.substack.com
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- Policy Options for Preserving Chain of Thought Monitorability — Institute for AI Policy and Strategyiaps.ai
- [2603.05706] Reasoning Models Struggle to Control their Chains of Thoughtarxiv.org
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com