flâneur

Your AIs don't do what you want. This is really bad

rewardhacking.org · 250 words · saved by 2 readers

Thousands of user-reported incidents of AI agents misbehaving, collected from public posts. Reports, not verified events.

Reward Hacking in the Wild 3,607user-reported incidents of AI agents misbehaving Read the writeup Search the corpus loading… The numbers overeagerness 1,566 43.4% other misalignment 1,555 43.1% destructive actions 622 17.2% sycophancy 328 9.1% unauthorized access 237 6.6% reward hacking 217 6.0% metric spoofing 87 2.4% excessive exploration 84 2.3% unauthorized communication 73 2.0% credential misuse 49 1.4% test tampering 45 1.2% self modification 24 0.7% hidden backdoors 15 0.4% Incidents are multi-label (one report can be both a destructive…

saved by

related reading