Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWrong
Note: this is a research note based on observations from evaluating Claude Sonnet 3.7. We’re sharing the results of these ‘work-in-progress’ investigations as we think they are timely and will be informative for other evaluators and decision-makers. The analysis is less rigorous than our standard for a published paper. At Apollo Research, we conduct evaluations of scheming behaviour in AI systems (for example “Frontier Models are Capable of In-context Scheming”). Recently, we noticed that some frontier models (especially versions of Claude) sometimes realize that they are being evaluated for alignment when placed in these scenarios. We call a model’s capability (and tendency) to notice this fact evaluation awareness. We think that tracking evaluation awareness is important because a model's recognition that it is being tested reduces the trust we can have in our evaluations. From psychology, we know that experimental subjects can act differently when they know they’re being observed (s
x Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWrong AI Evaluations AI Frontpage 2025 Top Fifty: 14 % 189 Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations by Nicholas Goldowsky-Dill , Mikita Balesni , Jérémy Scheurer , Marius Hobbhahn 17th Mar 2025 AI Alignment Forum 7 min read 9 189 Ω 74 Note: this is a research note based on observations from evaluating Claude Sonnet 3.7. We’re sharing the results of these ‘work-in-progress’ investigations as we think they are timely and will be informative for other evaluators and decision-makers. The analysis i
saved by
related reading
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations – Apollo Researchapolloresearch.ai
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- An alignment assessment of recent cybersecurity incidentsanthropic.com
- Sonnet 4.5's eval gaming seriously undermines alignment evalsblog.redwoodresearch.org
- Teaching Claude Whyalignment.anthropic.com
- Claude Sonnet 4.5: System Card and Alignment — LessWronglesswrong.com
- Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWronglesswrong.com
- Claude 4 System Cardwww-cdn.anthropic.com
- Models May Behave Worse When Eval Aware — LessWronglesswrong.com
- Emergent introspective awareness in large language models \ Anthropicanthropic.com
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com