Cyber Competitions
Throughout 2025, we have been quietly entering Claude in cybersecurity competitions designed primarily for humans. Now, we want to share what we have learned. In many of these competitions Claude did pretty well, often placing in the top 25% of competitors. However, it lagged behind the best human teams at the toughest challenges. Our experience testing Claude in cyber competitions highlights the potential for AI to alter the offense-defense balance by making it easier for attackers to automate the exploitation of basic vulnerabilities. More research and development into AI-enabled cyber defense and resilience is needed to counter this development. AI is poised to transform the domain of cybersecurity. Anthropic’s Safeguards team recently identified and banned a user with limited coding abilities leveraging Claude to develop malware. Research suggests that this lowering of the bar for expertise needed to pose a threat, combined with the falling costs of large language models (LLMs), pr
Frontier Red Team Claude is competitive with humans in (some) cyber competitions Aug 9, 2025 Throughout 2025, we have been quietly entering Claude in cybersecurity competitions designed primarily for humans. Now, we want to share what we have learned. In many of these competitions Claude did pretty well, often placing in the top 25% of competitors. However, it lagged behind the best human teams at the toughest challenges. Our experience testing Claude in cyber competitions highlights the potential for AI to alter the offense-defense balance by making it easier for attackers to automate the exp
Explore this link on the map →related reading
- Disrupting the first reported AI-orchestrated cyber espionage campaign \ Anthropicanthropic.com
- When AI builds itself \ Anthropicanthropic.com
- Project Glasswing: Securing critical software for the AI era \ Anthropicanthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Cybersecurity Looks Like Proof of Work Nowdbreunig.com
- Claude Fable 5 and Claude Mythos 5 \ Anthropicanthropic.com
- LLM-discovered 0 days \ Anthropicred.anthropic.com
- How we contain Claude across products \ Anthropicanthropic.com
- Investigating three real-world incidents in our cybersecurity evaluations \ Anthropicanthropic.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- AI excels at code competitions, struggles with real workblog.peterwildeford.com