Anthropic repeatedly accidentally trained against the CoT, demonstrating inadequate processes
blog.redwoodresearch.org · 1,624 words · saved by 1 readers
Safely navigating the intelligence explosion will require much more careful development
It turns out that Anthropic accidentally trained against the chain of thought of Claude Mythos Preview in around 8% of training episodes. This is at least the second independent incident in which Anthropic accidentally exposed their model’s CoT to the oversight signal. In more powerful systems, this kind of failure would jeopardize safely navigating the intelligence explosion. It’s crucial to build good processes to ensure development is executed according to plan, especially as human oversight becomes spread thin over increasing amounts of potentially untrusted and sloppy AI labor. This…
saved by
related reading
- Anthropic repeatedly accidentally trained against the CoT, demonstrating inadequate processessubstack.com
- Teaching Claude Whyalignment.anthropic.com
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- [2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv.org
- Teaching Claude why \ Anthropicanthropic.com
- An alignment assessment of recent cybersecurity incidentsanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Improving our alignment and security practicesanthropic.com
- Claude Mythos knows when it's breaking the rules — and tries to hide itsubstack.com
- The Most Forbidden Technique — LessWronglesswrong.com