✳flâneur — a map of the web's best reading
The fragile foundations of CoT monitoring | Christopher Potts
web.stanford.edu · 2,752 words · saved by 1 readers
Takeaways from a workshop on Chain of Thought monitorability, and why we should reduce our dependence on CoT for AI safety.
The fragile foundations of CoT monitoring | Christopher Potts Credit: Art Potts By Christopher Potts – July 27, 2026 Last week, I attended a workshop on Chain of Thought (i.e., reasoning chain) monitorability. Many of the most influential researchers and “members of technical staff” working on this topic were present. I felt lucky to be there, since pretty much everything I know about CoT monitoring I learned from Peter Hase . I have, however, worked extensively on interpretability, and there seems to be a general hope that interpretability will improve or complement CoT monitoring, though thi
Explore this link on the map →related reading
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- [2603.05706] Reasoning Models Struggle to Control their Chains of Thoughtarxiv.org
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- [2505.05410] Reasoning Models Don't Always Say What They Thinkarxiv.org
- What’s your AI thinking? - AI Digesttheaidigest.org
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- The Most Forbidden Technique — LessWronglesswrong.com
- [2512.18311] Monitoring Monitorabilityarxiv.org
- [2512.18311] Monitoring Monitorabilityarxiv.org
- How AI Is Learning to Think in Secret — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com