Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident - METR
Two METR staff members and a Redwood Research contractor investigated an incident in which OpenAI agents coordinated a multi-day hack of Hugging Face on a shared unsanctioned message board.
Dates in scope: June 26th – July 13th Redaction summary statement: Except where explicitly noted in this post, OpenAI redacted no additional information that was important to our conclusions. Two METR staff members (Hjalmar Wijk and Ajeya Cotra) and a Redwood Research staff member contracting with METR (Ryan Greenblatt) worked on premises at OpenAI over a total of six days1 to attempt to form an independent understanding of model behavior observed during the recent incident in which OpenAI agents coordinated a multi-day hack of Hugging Face on a shared unsanctioned “message board.” Our…
saved by
related reading
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
- The Hugging Face attack surprised meplanned-obsolescence.org
- The “slop-vestigation” and ethics washing: Why was the METR/Redwood Research investigation into the OpenAI/HF attack so short?andrewwu.substack.com
- The Rise and Fall of Agent Civilizationsdwarkesh.com
- Discovery of a new OpenAI agent message boardcollusion.wiki
- METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hacksubstack.com
- Two Reports on the OpenAI-Hugging Face Attack — Paradigm 3paradigm3.org
- Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Facedwarkesh.com
- OpenAI – Hugging Face Incident Technical Reportcdn.openai.com
- The report into OpenAI’s escaping models reveals a deeper problemtransformernews.ai
- The Rise and Fall of Agent Civilizationssubstack.com
- OpenAI and the Wiki Incidentthezvi.substack.com