flâneur — a map of the web's best reading

Reproducing steering against evaluation awareness in a large open-weight model — LessWrong

lesswrong.com · 7,820 words · saved by 1 readers

Produced as part of the UK AISI Model Transparency Team. Our team works on ensuring models don't subvert safety assessments, e.g. through evaluation…

x Reproducing steering against evaluation awareness in a large open-weight model — LessWrong Interpretability (ML & AI) AI Frontpage 90 Reproducing steering against evaluation awareness in a large open-weight model by Thomas Read , Bronson Schoen , Santiago Aranguri , Joseph Bloom 10th Apr 2026 18 min read 17 90 Produced as part of the UK AISI Model Transparency Team. Our team works on ensuring models don't subvert safety assessments, e.g. through evaluation awareness, sandbagging, or opaque reasoning. TL;DR We replicate Anthropic’s approach to using steering vectors to suppress evaluation awa

Explore this link on the map →

related reading