✳flâneur — a map of the web's best reading
Reproducing steering against evaluation awareness in a large open-weight model — LessWrong
lesswrong.com · 7,820 words · saved by 1 readers
Produced as part of the UK AISI Model Transparency Team. Our team works on ensuring models don't subvert safety assessments, e.g. through evaluation…
x Reproducing steering against evaluation awareness in a large open-weight model — LessWrong Interpretability (ML & AI) AI Frontpage 90 Reproducing steering against evaluation awareness in a large open-weight model by Thomas Read , Bronson Schoen , Santiago Aranguri , Joseph Bloom 10th Apr 2026 18 min read 17 90 Produced as part of the UK AISI Model Transparency Team. Our team works on ensuring models don't subvert safety assessments, e.g. through evaluation awareness, sandbagging, or opaque reasoning. TL;DR We replicate Anthropic’s approach to using steering vectors to suppress evaluation awa
Explore this link on the map →related reading
- Steering Might Stop Working Soon — LessWronglesswrong.com
- Activation Steering in 2026: A Practitioner's Field Guide | Subhadip Mitrasubhadipmitra.com
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWronglesswrong.com
- Realistic Evaluations Will Not Prevent Evaluation Awareness — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Models May Behave Worse When Eval Aware — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Teaching Models to Dream of Better Monitors through Evaluation Conditioned Training — LessWronglesswrong.com
- Metagaming matters for training, evaluation, and oversightalignment.openai.com
- Sonnet 4.5's eval gaming seriously undermines alignment evalsblog.redwoodresearch.org
- How far does alignment midtraining generalize?alignment.openai.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org