Your AIs don't do what you want. This is really bad
rewardhacking.org · 250 words · saved by 2 readers
Thousands of user-reported incidents of AI agents misbehaving, collected from public posts. Reports, not verified events.
Reward Hacking in the Wild 3,607user-reported incidents of AI agents misbehaving Read the writeup Search the corpus loading… The numbers overeagerness 1,566 43.4% other misalignment 1,555 43.1% destructive actions 622 17.2% sycophancy 328 9.1% unauthorized access 237 6.6% reward hacking 217 6.0% metric spoofing 87 2.4% excessive exploration 84 2.3% unauthorized communication 73 2.0% credential misuse 49 1.4% test tampering 45 1.2% self modification 24 0.7% hidden backdoors 15 0.4% Incidents are multi-label (one report can be both a destructive…
saved by
related reading
- METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hacksubstack.com
- Your AIs don't do what you want. This is really badreward-hacking-in-the-wild.vercel.app
- Your AIs don't do what you want. This is really badrewardhacking.org
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI Workaisi.gov.uk
- Why are AI agents lying, cheating and coordinating?yoshuabengio.org
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Countering misuse of AI: September 2026 / Anthropicanthropic.com
- Rogue AI Trackerrogueaitracker.com