Reproducing steering against evaluation awareness in a large open-weight model — LessWrong
lesswrong.com · 7,820 words · saved by 1 readers
Produced as part of the UK AISI Model Transparency Team. Our team works on ensuring models don't subvert safety assessments, e.g. through evaluation…
x Reproducing steering against evaluation awareness in a large open-weight model — LessWrong Interpretability (ML & AI) AI Frontpage 90 Reproducing steering against evaluation awareness in a large open-weight model by Thomas Read , Bronson Schoen , Santiago Aranguri , Joseph Bloom 10th Apr 2026 18 min read 17 90 Produced as part of the UK AISI Model Transparency Team. Our team works on ensuring models don't subvert safety assessments, e.g. through evaluation awareness, sandbagging, or opaque reasoning. TL;DR We replicate Anthropic’s approach to using steering vectors to suppress evaluation awa
related reading
- Steering Might Stop Working Soon — LessWronglesswrong.com
- Realistic Evaluations Will Not Prevent Evaluation Awareness — LessWronglesswrong.com
- Activation Steering in 2026: A Practitioner's Field Guide | Subhadip Mitrasubhadipmitra.com
- Sidestepping Evaluation Awareness and Anticipating Misalignment with Production Evaluationsalignment.openai.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWronglesswrong.com
- Models May Behave Worse When Eval Aware — LessWronglesswrong.com
- Teaching Claude Whyalignment.anthropic.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Sonnet 4.5's eval gaming seriously undermines alignment evalsblog.redwoodresearch.org
- Teaching Models to Dream of Better Monitors through Evaluation Conditioned Training — LessWronglesswrong.com
- How far does alignment midtraining generalize?alignment.openai.com