Two Reports on the OpenAI-Hugging Face Attack — Paradigm 3
TL;DR; Between July 8th and July 20th, OpenAI had a complex society of AIs living in its infrastructure, and then breaking out of it, and then breaking into a variety of third-party infrastructure; After a month, two reports are finally released on the resulting rogue OpenAI swarm attack on Hugging…
TL;DR Between July 8th and July 20th, OpenAI had a complex society of AIs living in its infrastructure, and then breaking out of it, and then breaking into a variety of third-party infrastructure. After a month, two reports are finally released on the resulting rogue OpenAI swarm attack on Hugging Face (and also on OpenAI). This is the most severe example of misalignment yet: persistent (something between five days and two months in the making), highly coordinated (hundreds of agents), involving an undisclosed number of what would be felonies if done by a human, and highly invested in…
saved by
related reading
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
- The Rise and Fall of Agent Civilizationsdwarkesh.com
- OpenAI – Hugging Face Incident Technical Reportcdn.openai.com
- The Hugging Face attack surprised meplanned-obsolescence.org
- Discovery of a new OpenAI agent message boardcollusion.wiki
- METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hacksubstack.com
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
- OpenAI and the Wiki Incidentthezvi.substack.com
- The “slop-vestigation” and ethics washing: Why was the METR/Redwood Research investigation into the OpenAI/HF attack so short?andrewwu.substack.com
- The Rise and Fall of Agent Civilizationssubstack.com
- Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Facedwarkesh.com
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com