✳flâneur — a map of the web's best reading
Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWrong
lesswrong.com · 12,139 words · saved by 1 readers
This post is written in our personal capacity. …
x Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWrong AI Evaluations AI Frontpage 46 Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face by Tim Hua , aditya singh 3rd Aug 2026 AI Alignment Forum 44 min read 2 46 Ω 19 This post is written in our personal capacity. Three-Minute Executive Summary An OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation. In this post, we describe the ambitious, comprehensive alignment evaluation we would run on this mode
Explore this link on the map →related reading
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- Natural emergent misalignment from reward hacking \ Anthropicanthropic.com
- Your AIs don't do what you want. This is really badrewardhacking.org
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org