✳flâneur — a map of the web's best reading
Teaching Models to Dream of Better Monitors through Evaluation Conditioned Training — LessWrong
lesswrong.com · 4,288 words · saved by 1 readers
Edit: we renamed the technique from Monitor Sensitive Training (MST) to Evaluation Conditioned Training (ECT)[1] …
x Teaching Models to Dream of Better Monitors through Evaluation Conditioned Training — LessWrong Outer Alignment Scalable Oversight AI Frontpage 49 Teaching Models to Dream of Better Monitors through Evaluation Conditioned Training by Alec Harris , Kasey C , Archie Chaudhury , yix 19th Mar 2026 12 min read 2 49 Edit: we renamed the technique from Monitor Sensitive Training (MST) to Evaluation Conditioned Training (ECT) [1] TL;DR We introduce Evaluation Conditioned Training (ECT), a new post-training technique where we augment training data with evaluation labels that describe how evaluation i
Explore this link on the map →saved by
related reading
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Teaching Claude why \ Anthropicanthropic.com
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- We need a better way to evaluate emergent misalignment — LessWronglesswrong.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- How far does alignment midtraining generalize?alignment.openai.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Reproducing steering against evaluation awareness in a large open-weight model — LessWronglesswrong.com
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWronglesswrong.com
- Unsupervised Elicitationalignment.anthropic.com