Should We Train Against (CoT) Monitors? — LessWrong
The question I actually try to answer in this post is a broader one (that doesn't work as well as a title): Should we incorporate proxies for desired…
x Should We Train Against (CoT) Monitors? — LessWrong Aether AI Frontpage 50 Should We Train Against (CoT) Monitors? by RohanS 23rd Apr 2026 39 min read 7 50 The question I actually try to answer in this post is a broader one (that doesn't work as well as a title): Should we incorporate proxies for desired behavior into LLM alignment training? Epistemic status: My best guess. I tentatively claim that we should be more open to incorporating proxies for desired behavior into LLM training, but I try to clarify the spectrum of possible answers beyond just 'yes' and 'no,' and I try to present and s
saved by
related reading
- Alignment Faking Mitigationsalignment.anthropic.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- [2512.11949] Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitorsarxiv.org
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Why We Are Excited About Confessionsalignment.openai.com
- Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimesarxiv.org
- We need 3rd party Training-Run Assessments — LessWronglesswrong.com
- The Most Forbidden Technique — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org