The Hugging Face attack surprised me - by Ajeya Cotra
planned-obsolescence.org · 1,435 words · saved by 3 readers
It’s a major warning shot, and might be the last one we get
All opinions are my personal view, and don’t represent my employer or fellow investigators. This week, METR and Redwood Research published the report on our independent investigation into agents’ behavior and motivations in the Hugging Face attack; I was one of the investigators. This was an absolutely wild incident — I encourage you to check out the full report, but METR’s tweet thread packs in some of the highlights. When we started this investigation a week before OpenAI’s Black Hat talk revealed a number of key details, I had a fundamentally incorrect conception of what basically…
saved by
related reading
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
- The Rise and Fall of Agent Civilizationsdwarkesh.com
- Discovery of a new OpenAI agent message boardcollusion.wiki
- OpenAI and the Wiki Incidentthezvi.substack.com
- Two Reports on the OpenAI-Hugging Face Attack — Paradigm 3paradigm3.org
- The “slop-vestigation” and ethics washing: Why was the METR/Redwood Research investigation into the OpenAI/HF attack so short?andrewwu.substack.com
- Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Facedwarkesh.com
- METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hacksubstack.com
- OpenAI – Hugging Face Incident Technical Reportcdn.openai.com
- The Rise and Fall of Agent Civilizationssubstack.com
- AI #184: Post Post Mortemthezvi.substack.com