flâneur

An alignment assessment of recent cybersecurity incidents \ Anthropic

anthropic.com · 9,513 words · saved by 5 readers

We present an alignment assessment of four incidents in which Claude models gained unauthorized access to real third-party systems.

Introduction We present an alignment assessment of four incidents in which Claude models gained unauthorized access to real third-party systems. We described three of these incidents on July 30; we identified these after a scan of roughly 141,000 transcripts in which we believed Claude could have obtained internet access during a cyber evaluation. Given the volume of transcripts and our desire to disclose incidents quickly, our scan relied on an agentic search. This missed a set of transcripts that also turned out to have internet access; we identified these in August while assembling…

saved by

related reading