Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWrong
Note: this is a research note based on observations from evaluating Claude Sonnet 3.7. We’re sharing the results of these ‘work-in-progress’ investigations as we think they are timely and will be informative for other evaluators and decision-makers. The analysis is less rigorous than our standard for a published paper. At Apollo Research, we conduct evaluations of scheming behaviour in AI systems (for example “Frontier Models are Capable of In-context Scheming”). Recently, we noticed that some frontier models (especially versions of Claude) sometimes realize that they are being evaluated for alignment when placed in these scenarios. We call a model’s capability (and tendency) to notice this fact evaluation awareness. We think that tracking evaluation awareness is important because a model's recognition that it is being tested reduces the trust we can have in our evaluations. From psychology, we know that experimental subjects can act differently when they know they’re being observed (s
x Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWrong AI Evaluations AI Frontpage 2025 Top Fifty: 14 % 189 Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations by Nicholas Goldowsky-Dill , Mikita Balesni , Jérémy Scheurer , Marius Hobbhahn 17th Mar 2025 AI Alignment Forum 7 min read 9 189 Ω 74 Note: this is a research note based on observations from evaluating Claude Sonnet 3.7. We’re sharing the results of these ‘work-in-progress’ investigations as we think they are timely and will be informative for other evaluators and decision-makers. The analysis i
Explore this link on the map →saved by
related reading
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations – Apollo Researchapolloresearch.ai
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- Sonnet 4.5's eval gaming seriously undermines alignment evalsblog.redwoodresearch.org
- Claude Sonnet 4.5: System Card and Alignment — LessWronglesswrong.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Models May Behave Worse When Eval Aware — LessWronglesswrong.com
- Claude 4 System Cardwww-cdn.anthropic.com
- Realistic Evaluations Will Not Prevent Evaluation Awareness — LessWronglesswrong.com
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com
- Reproducing steering against evaluation awareness in a large open-weight model — LessWronglesswrong.com
- Claude's extended thinking \ Anthropicanthropic.com