Misleading Metaphors, Real Risks - by Melanie Mitchell
aiguide.substack.com · 3,230 words · saved by 1 readers
What To Fear from AI Agents and How to Reclaim Our Human Agency
The recent clamor around AI safety has been deafening. Wired reported that during an evaluation of their AI systems, OpenAI “lost control of two AI models.”1 Articles in the New York Times stated that “rogue [AI] agents had created their own message board to communicate with one another”2 and that even after OpenAI noticed problems and closed some security holes, “the swarm broke out of its cage again using hacks that were heretofore undiscovered by humans. This time, the swarm ran free for about a week before it was noticed.”3 Possibly in response to this and other AI agent hacking…
saved by
related reading
- The Rise and Fall of Agent Civilizationsdwarkesh.com
- Why are AI agents lying, cheating and coordinating?yoshuabengio.org
- AI 2027ai-2027.com
- Can A.I. “Go Rogue”?newyorker.com
- The Rise and Fall of Agent Civilizationssubstack.com
- Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Facedwarkesh.com
- AI 2027ai-2027.com
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI Workaisi.gov.uk
- METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hacksubstack.com
- Walter Mitty Effects in AI Incident Analysescontraptions.venkateshrao.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com