[2603.11214] Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios
Abstract:We evaluate the autonomous cyber-attack capabilities of frontier AI models on two purpose-built cyber ranges-a 32-step corporate network attack and a 7-step industrial control system attack-that require chaining heterogeneous capabilities across extended action sequences. By comparing seven models released over an eighteen-month period (August 2024 to February 2026) at varying inference-time compute budgets, we observe two capability trends. First, model performance scales log-linearly with inference-time compute, with no observed plateau-increasing from 10M to 100M tokens yields gains of up to 59%, requiring no specific technical sophistication from the operator. Second, each successive model generation outperforms its predecessor at fixed token budgets: on the corporate network range, average steps completed at 10M tokens rose from 1.7 (GPT-4o, August 2024) to 9.8 (Opus 4.6, February 2026). The best single run completed 22 of 32 steps, corresponding to roughly 6 of the estimated 14 hours a human expert would need. On the industrial control system range, performance remains limited, though the most recent models are the first to reliably complete steps, averaging 1.2-1.4 of 7 (max 3).
# link_n76n5j0srf.pdf ## Metadata - PDFFormatVersion=1.7 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Linus Folkerts; Will Payne; Simon Inman; Philippos Giavridis; Joe Skinner; Sam Deverett; James Aung; Ekin Zorer; Michael Schmatz; Mahmoud Ghanem; John Wilkinson; Alan Steer; Vy Hong; Jessica Wang - Creator=arXiv GenPDF (tex2pdf:8e18bb9) - Custom.DOI=https://doi.org/10.48550/arXiv.2603.11214 - Custom.License=http://creativecommons.org/licenses/by/4.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592
Explore this link on the map →saved by
related reading
- Disrupting the first reported AI-orchestrated cyber espionage campaign \ Anthropicanthropic.com
- AI 2027ai-2027.com
- My picture of the present in AI — LessWronglesswrong.com
- Offense at Scale: How Frontier AI Lowers the Cost of Cyber Attacks - Irregularirregular.com
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- AI in 2025: gestalt — LessWronglesswrong.com
- AI 2027ai-2027.com
- How fast is AI improving? - AI Digesttheaidigest.org
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Cybersecurity Looks Like Proof of Work Nowdbreunig.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Emergent Cyber Behavior: When AI Agents Become Offensive Threat Actors - Irregularirregular.com