Training fails to elicit subtle reasoning in current language models
While recent AI systems achieve strong performance through human-readable reasoning that should be simple to monitor (OpenAI, 2024, Anthropic, 2025), we investigate whether models can learn to reason about malicious side tasks while making that reasoning appear benign. We find that Sonnet 3.7 can learn to evade either a reasoning monitor, by persuading the monitor that a blatant backdoor is benign, or an output-only monitor, by devising sophisticated backdoors that the output-only monitor doesn’t detect. But when trained to evade both reasoning and output-only monitors, Sonnet 3.7 is unable to use reasoning to improve its backdoor success rate without triggering a reasoning monitor. Like previous work (Baker et al., 2025, Emmons et al., 2025), our results suggest that reasoning monitors can provide strong assurance that language models are not pursuing reasoning-heavy malign side tasks, but that additional mitigations may be required for robustness to monitor persuasion. As models beco
Training fails to elicit subtle reasoning in current language models Alignment Science Blog Training fails to elicit subtle reasoning in current language models Training fails to elicit subtle reasoning in current language models tl;dr While recent AI systems achieve strong performance through human-readable reasoning that should be simple to monitor ( OpenAI, 2024 , Anthropic, 2025 ), we investigate whether models can learn to reason about malicious side tasks while making that reasoning appear benign. We find that Sonnet 3.7 can learn to evade either a reasoning monitor, by persuading the mo
Explore this link on the map →related reading
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- DeepSeek-R1arxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- How AI Is Learning to Think in Secret — LessWronglesswrong.com
- [2510.27338] Reasoning Models Sometimes Output Illegible Chains of Thoughtarxiv.org
- [2510.09714] All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Languagearxiv.org
- 2312.06942arxiv.org