Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
dwarkesh.com · 10,306 words · saved by 1 readers
"This might be the clearest warning shot we ever get."
Ajeya Cotra is a researcher at METR, where she works on threat modeling for loss-of-control risks from advanced AI. Before that, she led the technical AI safety program at what is now Coefficient Giving. She is one the three authors of METR and Redwood Research’s “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident”. We go through not only what she and her coauthors discovered during this investigation, but what it means for how we should train future, smarter AIs which might be involved in the process of recursive…
saved by
related reading
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
- The Rise and Fall of Agent Civilizationsdwarkesh.com
- The Hugging Face attack surprised meplanned-obsolescence.org
- The “slop-vestigation” and ethics washing: Why was the METR/Redwood Research investigation into the OpenAI/HF attack so short?andrewwu.substack.com
- METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hacksubstack.com
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
- AI 2027ai-2027.com
- Two Reports on the OpenAI-Hugging Face Attack — Paradigm 3paradigm3.org
- Discovery of a new OpenAI agent message boardcollusion.wiki
- The Rise and Fall of Agent Civilizationssubstack.com
- OpenAI – Hugging Face Incident Technical Reportcdn.openai.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com