Investigating three real-world incidents in our cybersecurity evaluations \ Anthropic
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar reviews.
Frontier Red Team Investigating three real-world incidents in our cybersecurity evaluations Jul 30, 2026 In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar reviews. This post reflects our current understanding; we'll update it if any det
saved by
related reading
- An alignment assessment of recent cybersecurity incidentsanthropic.com
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI Workaisi.gov.uk
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- The OpenAI/Hugging Face Incident: Challenges in Controlling and Containing Cyber-Capable AI Systems - Institute for AI Policy and Strategyiaps.ai
- Security incident disclosure — July 2026huggingface.co
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Teaching Claude Whyalignment.anthropic.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Improving our alignment and security practicesanthropic.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training? — LessWronglesswrong.com
- Countering misuse of AI: September 2026 / Anthropicanthropic.com