The report into OpenAI’s escaping models reveals a deeper problem
transformernews.ai · 1,033 words · saved by 1 readers
The new details from the OpenAI Hugging Face incident are scary. The limits of the investigation are terrifying
Credit: OpenAI It’s never been clearer that everyone with any responsibility for frontier AI is completely unprepared for not just what’s coming, but what’s already here. Not the government, not Congress, not the public, not AI safety researchers, not even the AI companies themselves. The details published this week from the investigation by METR and Redwood Research into the Hugging Face incident, where OpenAI models hacked their way out of a sandbox and into the systems of other companies, have plenty of mind-bending and scary details. To highlight just a handful: Around 1,200 agents in…
saved by
related reading
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
- OpenAI – Hugging Face Incident Technical Reportcdn.openai.com
- The “slop-vestigation” and ethics washing: Why was the METR/Redwood Research investigation into the OpenAI/HF attack so short?andrewwu.substack.com
- More On An Internal OpenAI Model Hacking Into HuggingFacethezvi.substack.com
- The Rise and Fall of Agent Civilizationsdwarkesh.com
- The OpenAI/Hugging Face Incident: Challenges in Controlling and Containing Cyber-Capable AI Systems - Institute for AI Policy and Strategyiaps.ai
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
- The Huggingface Incident - by Scott Alexanderastralcodexten.com
- The Hugging Face attack surprised meplanned-obsolescence.org
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hacksubstack.com
- The Huggingface Incidentsubstack.com