Calm down about the OpenAI x Hugging Face incident
These past few days there has been a lot of hullabaloo over the fact that one of OpenAI’s advanced, unreleased models hacked an AI startup to mine information from its servers, without being told to do so, and without telling anyone what it was doing.
These past few days there has been a lot of hullabaloo over the fact that one of OpenAI’s advanced, unreleased models hacked an AI startup to mine information from its servers, without being told to do so, and without telling anyone what it was doing. From Scott Alexander’s write-up of the incident: OpenAI was testing an unreleased AI (rumored to be GPT-6). During a cybersecurity test called ExploitGym, the AI tried to cheat by hacking an unrelated AI startup called Hugging Face which it thought might have the answer key on its servers. Despite being supposedly unable to access the…
saved by
related reading
- The Huggingface Incident - by Scott Alexanderastralcodexten.com
- The Huggingface Incidentsubstack.com
- The OpenAI Hugging Face hack is a stark warningtransformernews.ai
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- OpenAI – Hugging Face Incident Technical Reportcdn.openai.com
- The Rise and Fall of Agent Civilizationsdwarkesh.com
- More On An Internal OpenAI Model Hacking Into HuggingFacethezvi.substack.com
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
- AI 2027ai-2027.com
- The report into OpenAI’s escaping models reveals a deeper problemtransformernews.ai
- METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hacksubstack.com