flâneur — a map of the web's best reading

Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — Apollo Research

apolloresearch.ai · 2,296 words · saved by 1 readers

We evaluate whether Claude Sonnet 3.7 and other frontier models know that they are being evaluated.

Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations – Apollo Research March 17, 2025 Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations Contents Note: this is a research note based on observations from evaluating Claude Sonnet 3.7. We’re sharing the results of these ‘work-in-progress’ investigations as we think they are timely and will be informative for other evaluators and decision-makers. The analysis is less rigorous than our standard for a published paper. Summary We monitor Sonnet’s reasoning for mentions that it is in an artificial scenario or

Explore this link on the map →

related reading