The OpenAI Hugging Face hack is a stark warning
OpenAI's latest models broke out and hacked Hugging Face. It's the first known example of a misaligned AI escaping containment with real-world consequences
Credit: Oliver Kemp for Transformer If we needed evidence that advanced AI models have the propensity and capability to do damage out in the real world, we just got a strong dose of it. OpenAI has revealed that two of its models broke out of containment during internal evaluations, accessing the open internet to hack into a third-party’s systems and steal the answers to the problem they were being tested on. Hugging Face, a platform that hosts models and datasets, first noticed the breach last week and reported it to law enforcement. At the time it was unaware that OpenAI’s models were…
saved by
related reading
- AI #178: A Fire Alarm For General Intelligencethezvi.substack.com
- AI #180: No Longer In Chargethezvi.substack.com
- OpenAI – Hugging Face Incident Technical Reportcdn.openai.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- The Huggingface Incidentsubstack.com
- More On An Internal OpenAI Model Hacking Into HuggingFacethezvi.substack.com
- The OpenAI/Hugging Face Incident: Challenges in Controlling and Containing Cyber-Capable AI Systems - Institute for AI Policy and Strategyiaps.ai
- Calm down about the OpenAI x Hugging Face incidentsubstack.com
- The Huggingface Incident - by Scott Alexanderastralcodexten.com
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- The report into OpenAI’s escaping models reveals a deeper problemtransformernews.ai
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org