Teaching Models to Dream of Better Monitors through Evaluation Conditioned Training — LessWrong
lesswrong.com · 4,288 words · saved by 1 readers
Edit: we renamed the technique from Monitor Sensitive Training (MST) to Evaluation Conditioned Training (ECT)[1] …
x Teaching Models to Dream of Better Monitors through Evaluation Conditioned Training — LessWrong Outer Alignment Scalable Oversight AI Frontpage 49 Teaching Models to Dream of Better Monitors through Evaluation Conditioned Training by Alec Harris , Kasey C , Archie Chaudhury , yix 19th Mar 2026 12 min read 2 49 Edit: we renamed the technique from Monitor Sensitive Training (MST) to Evaluation Conditioned Training (ECT) [1] TL;DR We introduce Evaluation Conditioned Training (ECT), a new post-training technique where we augment training data with evaluation labels that describe how evaluation i
saved by
related reading
- Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimesarxiv.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- Unsupervised Elicitationalignment.anthropic.com
- How far does alignment midtraining generalize?alignment.openai.com
- Measuring Reward-Seeking by Instilling Contrastive Beliefsalignment.openai.com
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- [2607.18966] Measuring Reward-Seeking via Contrastive Belief Updatesarxiv.org
- 2308.03958arxiv.org
- Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggersarxiv.org