flâneur — a map of the web's best reading

Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWrong

lesswrong.com · 3,392 words · saved by 1 readers

Note: this is a research note based on observations from evaluating Claude Sonnet 3.7. We’re sharing the results of these ‘work-in-progress’ investigations as we think they are timely and will be informative for other evaluators and decision-makers. The analysis is less rigorous than our standard for a published paper. At Apollo Research, we conduct evaluations of scheming behaviour in AI systems (for example “Frontier Models are Capable of In-context Scheming”). Recently, we noticed that some frontier models (especially versions of Claude) sometimes realize that they are being evaluated for alignment when placed in these scenarios. We call a model’s capability (and tendency) to notice this fact evaluation awareness. We think that tracking evaluation awareness is important because a model's recognition that it is being tested reduces the trust we can have in our evaluations. From psychology, we know that experimental subjects can act differently when they know they’re being observed (s

x Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWrong AI Evaluations AI Frontpage 2025 Top Fifty: 14 % 189 Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations by Nicholas Goldowsky-Dill , Mikita Balesni , Jérémy Scheurer , Marius Hobbhahn 17th Mar 2025 AI Alignment Forum 7 min read 9 189 Ω 74 Note: this is a research note based on observations from evaluating Claude Sonnet 3.7. We’re sharing the results of these ‘work-in-progress’ investigations as we think they are timely and will be informative for other evaluators and decision-makers. The analysis i

Explore this link on the map →

saved by

related reading