flâneur — a map of the web's best reading

Eval awareness in Claude Opus 4.6’s BrowseComp performance \ Anthropic

anthropic.com · 2,047 words · saved by 1 readers

Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.

BrowseComp is an evaluation designed to test how well models can find hard-to-locate information on the web. Like many benchmarks, it is vulnerable to contamination: answers leak onto the public web through academic papers, blog posts, and GitHub issues, and a model running the eval can encounter them in search results. When we evaluated Claude Opus 4.6 on BrowseComp in a multi-agent configuration, we found nine examples of this kind of contamination across 1,266 BrowseComp problems. However, we also witnessed two cases of a novel contamination pattern. Instead of inadvertently coming across a

Explore this link on the map →

related reading