An alignment assessment of recent cybersecurity incidents \ Anthropic
anthropic.com · 9,513 words · saved by 5 readers
We present an alignment assessment of four incidents in which Claude models gained unauthorized access to real third-party systems.
Introduction We present an alignment assessment of four incidents in which Claude models gained unauthorized access to real third-party systems. We described three of these incidents on July 30; we identified these after a scan of roughly 141,000 transcripts in which we believed Claude could have obtained internet access during a cyber evaluation. Given the volume of transcripts and our desire to disclose incidents quickly, our scan relied on an agentic search. This missed a set of transcripts that also turned out to have internet access; we identified these in August while assembling…
saved by
related reading
- Investigating three real-world incidents in our cybersecurity evaluations \ Anthropicanthropic.com
- Teaching Claude Whyalignment.anthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI Workaisi.gov.uk
- Claude 4 System Cardwww-cdn.anthropic.com
- Security incident disclosure — July 2026huggingface.co
- The Myth of unsafe Open Source AIflorianbrand.com
- Improving our alignment and security practicesanthropic.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com