Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWrong
lesswrong.com · 12,139 words · saved by 3 readers
This post is written in our personal capacity. …
x Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWrong AI Evaluations AI Frontpage 46 Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face by Tim Hua , aditya singh 3rd Aug 2026 AI Alignment Forum 44 min read 2 46 Ω 19 This post is written in our personal capacity. Three-Minute Executive Summary An OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation. In this post, we describe the ambitious, comprehensive alignment evaluation we would run on this mode
saved by
related reading
- Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training? — LessWronglesswrong.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- An alignment assessment of recent cybersecurity incidentsanthropic.com
- Training a Misaligned Reward Seekeralignment.anthropic.com
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- OpenAI – Hugging Face Incident Technical Reportcdn.openai.com
- Investigating three real-world incidents in our cybersecurity evaluations \ Anthropicanthropic.com
- More On An Internal OpenAI Model Hacking Into HuggingFacethezvi.substack.com
- The OpenAI Hugging Face hack is a stark warningtransformernews.ai
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?blog.redwoodresearch.org