Should We Train Against (CoT) Monitors? — LessWrong
The question I actually try to answer in this post is a broader one (that doesn't work as well as a title): Should we incorporate proxies for desired…
x Should We Train Against (CoT) Monitors? — LessWrong Aether AI Frontpage 50 Should We Train Against (CoT) Monitors? by RohanS 23rd Apr 2026 39 min read 7 50 The question I actually try to answer in this post is a broader one (that doesn't work as well as a title): Should we incorporate proxies for desired behavior into LLM alignment training? Epistemic status: My best guess. I tentatively claim that we should be more open to incorporating proxies for desired behavior into LLM training, but I try to clarify the spectrum of possible answers beyond just 'yes' and 'no,' and I try to present and s
Explore this link on the map →saved by
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- The Most Forbidden Technique — LessWronglesswrong.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Alignment faking in large language modelsarxiv.org
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- How far does alignment midtraining generalize?alignment.openai.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Thomas Larsen's Shortform — LessWronglesswrong.com
- An AI alignment research agenda based on asymmetric debate and monitoring. — LessWronglesswrong.com