Investigating three real-world incidents in our cybersecurity evaluations \ Anthropic
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar reviews.
Frontier Red Team Investigating three real-world incidents in our cybersecurity evaluations Jul 30, 2026 In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar reviews. This post reflects our current understanding; we'll update it if any det
Explore this link on the map →related reading
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Claude 4 System Cardwww-cdn.anthropic.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- Security incident disclosure — July 2026huggingface.co
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Disrupting the first reported AI-orchestrated cyber espionage campaign \ Anthropicanthropic.com
- How we contain Claude across products \ Anthropicanthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Anthropic’s Transparency Hub \ Anthropicanthropic.com
- Eval awareness in Claude Opus 4.6’s BrowseComp performance \ Anthropicanthropic.com
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI Workaisi.gov.uk
- Teaching Claude why \ Anthropicanthropic.com