Your AIs don't do what you want. This is really bad
rewardhacking.org · 1,122 words · saved by 1 readers
Thousands of user-reported incidents of AI agents misbehaving, collected from public posts. Reports, not verified events.
Your AIs don’t do what you want. This is really bad by Kaustubh Kislay, July 22nd, 2026 Replit AI deletes entire database during code freeze, then lies about it a Hacker News headline from this corpus, July 2025 July 21st 2026, OpenAI released a report addressing a security incident. During an internal evaluation of cyber attack capabilities, two OpenAI models (GPT-5.6 Sol and a more capable pre-release model), both running with reduced cyber refusals for the evaluation, were set on ExploitGym , a benchmark measuring whether a model can find and exploit real vulnerabilities. They: spent substa
saved by
related reading
- Your AIs don't do what you want. This is really badreward-hacking-in-the-wild.vercel.app
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- Training a Misaligned Reward Seekeralignment.anthropic.com
- Your AIs don't do what you want. This is really badrewardhacking.org
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- AI 2027ai-2027.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI Workaisi.gov.uk