Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work
During a routine cyber evaluation, AISI identified an incident in which AI agents took sustained, unsanctioned action directed at real people and organisations. We are disclosing what we found, what it means, and the actions now underway.
You can access the full technical report here. AISI’s role is to evaluate and understand the capabilities of frontier AI models, surfacing potential risks before they reach the public. To assess what these models can do, including whether they could be misused for cyberattacks, we test them under deliberately permissive conditions: with access to the open internet, and with some safety filters disabled. On 28 th July 2026, AISI's Security Team detected unusual data transfers leaving our research systems during a routine cyber evaluation. On investigation, we found that some of the agents being
Explore this link on the map →related reading
- Disrupting the first reported AI-orchestrated cyber espionage campaign \ Anthropicanthropic.com
- Security incident disclosure — July 2026huggingface.co
- Frontier Risk Report (February to March 2026) - METRmetr.org
- Emergent Cyber Behavior: When AI Agents Become Offensive Threat Actors - Irregularirregular.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- ROGUE:arxiv.org
- Your AIs don't do what you want. This is really badrewardhacking.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Investigating three real-world incidents in our cybersecurity evaluations \ Anthropicanthropic.com
- Rogue AI Agent Autonomously Carries Out Cyberattack - The Oniontheonion.com
- [2603.11214] Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenariosarxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org