Autonomous hacking
I've tested an autonomous hacker on real products over the past month,[1] and it's very capable. Concretely, I could: 1. Hijack accounts at a major bank and log in as other users. 2. Bypass authorization in a major AI lab's product and access other users' private data, including uploaded
I've tested an autonomous hacker on real products over the past month, [1] and it's very capable. Concretely, I could: Hijack accounts at a major bank and log in as other users. Bypass authorization in a major AI lab's product and access other users' private data, including uploaded files. List all users of a big tech company's product and modify their files. Download the health records of anyone whose data was stored and sharable through a popular electronic health records system. [2] If I were malicious, I reasonably believe that I could cause billions of dollars in damage in less than a mon
Explore this link on the map →related reading
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Project Glasswing: Securing critical software for the AI era \ Anthropicanthropic.com
- AI 2027ai-2027.com
- Vulnerability Research Is Cooked - Quarrelsomesockpuppet.org
- Security incident disclosure — July 2026huggingface.co
- Disrupting the first reported AI-orchestrated cyber espionage campaign \ Anthropicanthropic.com
- Offense at Scale: How Frontier AI Lowers the Cost of Cyber Attacks - Irregularirregular.com
- Cybersecurity Looks Like Proof of Work Nowdbreunig.com
- The OpenAI/Hugging Face Incident: Challenges in Controlling and Containing Cyber-Capable AI Systems - Institute for AI Policy and Strategyiaps.ai
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Emergent Cyber Behavior: When AI Agents Become Offensive Threat Actors - Irregularirregular.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com