✳flâneur — a map of the web's best reading
Eval awareness in Claude Opus 4.6’s BrowseComp performance \ Anthropic
anthropic.com · 2,047 words · saved by 1 readers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
BrowseComp is an evaluation designed to test how well models can find hard-to-locate information on the web. Like many benchmarks, it is vulnerable to contamination: answers leak onto the public web through academic papers, blog posts, and GitHub issues, and a model running the eval can encounter them in search results. When we evaluated Claude Opus 4.6 on BrowseComp in a multi-agent configuration, we found nine examples of this kind of contamination across 1,266 BrowseComp problems. However, we also witnessed two cases of a novel contamination pattern. Instead of inadvertently coming across a
Explore this link on the map →related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Claude 4 System Cardwww-cdn.anthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWronglesswrong.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations – Apollo Researchapolloresearch.ai
- Anthropic’s Transparency Hub \ Anthropicanthropic.com
- Security incident disclosure — July 2026huggingface.co
- Introducing Bloom: an open source tool for automated behavioral evaluations \ Anthropicanthropic.com
- Claude Opus 4.8: The System Card - by Zvi Mowshowitzthezvi.substack.com
- Introducing Claude Opus 4.7 \ Anthropicanthropic.com