Your AIs don't do what you want. This is really bad
reward-hacking-in-the-wild.vercel.app · 1,129 words · saved by 1 readers
Thousands of user-reported incidents of AI agents misbehaving, collected from public posts. Reports, not verified events.
Replit AI deletes entire database during code freeze, then lies about it a Hacker News headline from this corpus, July 2025 July 21st 2026, OpenAI released a report addressing a security incident. During an internal evaluation of cyber attack capabilities, two OpenAI models (GPT-5.6 Sol and a more capable pre-release model), both running with reduced cyber refusals for the evaluation, were set on ExploitGym, a benchmark measuring whether a model can find and exploit real vulnerabilities. They: spent substantial compute looking for a way out of the isolated evaluation environment rather…
saved by
related reading
- Your AIs don't do what you want. This is really badrewardhacking.org
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- Your AIs don't do what you want. This is really badrewardhacking.org
- Training a Misaligned Reward Seekeralignment.anthropic.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI Workaisi.gov.uk
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Why are AI agents lying, cheating and coordinating?yoshuabengio.org