What’s your AI thinking? - AI Digest
AI is rapidly becoming more capable – the time horizon for coding tasks is doubling every 4-7 months. But we don’t actually know what these increasingly capable models are thinking. And that’s a problem. If we can’t tell what a model is thinking, then we can’t tell when it is downplaying its capabilities, cheating on tests, or straight up working against us. Luckily we do have a lead: the chain of thought (CoT). This CoT is used in all top-performing language models. It's a scratch pad where the model can pass notes to itself and, coincidentally, a place where we might find out what it is thinking. Except, the CoT isn’t always faithful. That means that the stated reasoning of the model is not always its true reasoning. And we are not sure yet how to improve that. However, some researchers now argue that we don’t need complete faithfulness. They argue monitorability is sufficient. While faithfulness means you can read the model’s mind and know what it is thinking. Monitorability means y
What’s your AI thinking? - AI Digest Demos and explainers About Join the team What’s your AI thinking? What is Chain of Thought? Chain of Thought Faithfulness Chain of Thought Monitorability Keeping the Chain of Thought Monitorable What’s your AI thinking? A step by step introduction to chain of thought monitorability August 4, 2025 Shoshannah Tekofsky Member of technical staff at AI Digest AI is rapidly becoming more capable – the time horizon for coding tasks is doubling every 4-7 months . But we don’t actually know what these increasingly capable models are thinking . And that’s a problem.
Explore this link on the map →related reading
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- How AI Is Learning to Think in Secret — LessWronglesswrong.com
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- [2603.05706] Reasoning Models Struggle to Control their Chains of Thoughtarxiv.org
- [2505.05410] Reasoning Models Don't Always Say What They Thinkarxiv.org
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- [2512.18311] Monitoring Monitorabilityarxiv.org
- [2512.18311] Monitoring Monitorabilityarxiv.org