Rogue AI Tracker
rogueaitracker.com · 1,330 words · saved by 1 readers
Rogue AI Tracker reviews public incidents and research to track observed autonomous AI-agent capabilities and critical milestones.
Milestones A swarm of AI agents runs cyberattack campaigns against multiple targets at once, exploiting vulnerabilities at scale, resulting in major security breaches across multiple victims. Long-horizon goal execution Long-horizon goal execution An agent carried out a long, multi-step task on its own, handling errors and decisions along the way. Current score: 8/10 Against live third parties, an agent pursuing an assigned goal independently chooses an unintended multi-step campaign, sustains it through changing state or failures, and achieves the objective it was working toward. OpenAI…
saved by
related reading
- CAIS AI Dashboarddashboard.safe.ai
- Lakera – Test your AI hacking skillsgandalf.lakera.ai
- Incidents | Rogue AI Trackerrogueaitracker.com
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI Workaisi.gov.uk
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Frontier Risk Report (February to March 2026) - METRmetr.org
- Security incident disclosure — July 2026huggingface.co
- There's An AI For That® — The front page of AItheresanaiforthat.com
- Donovan: AI Agents for the Public Sectorscale.com
- Is Agentic: AI Agent Readiness Score for your Site and Appis-agentic.com
- Frontier AI Cybersecurity Observatorycybergym.io
- Goodfire AIgoodfire.ai