✳flâneur — a map of the web's best reading
Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — Apollo Research
apolloresearch.ai · 2,296 words · saved by 1 readers
We evaluate whether Claude Sonnet 3.7 and other frontier models know that they are being evaluated.
Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations – Apollo Research March 17, 2025 Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations Contents Note: this is a research note based on observations from evaluating Claude Sonnet 3.7. We’re sharing the results of these ‘work-in-progress’ investigations as we think they are timely and will be informative for other evaluators and decision-makers. The analysis is less rigorous than our standard for a published paper. Summary We monitor Sonnet’s reasoning for mentions that it is in an artificial scenario or
Explore this link on the map →related reading
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWronglesswrong.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Claude Sonnet 4.5: System Card and Alignment — LessWronglesswrong.com
- Claude 4 System Cardwww-cdn.anthropic.com
- Sonnet 4.5's eval gaming seriously undermines alignment evalsblog.redwoodresearch.org
- Models May Behave Worse When Eval Aware — LessWronglesswrong.com
- Sandbagging with misaligned action - Chain-of-Thought Transcript - Anti-Schemingantischeming.ai
- Eval awareness in Claude Opus 4.6’s BrowseComp performance \ Anthropicanthropic.com
- Claude's extended thinking \ Anthropicanthropic.com
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com