flâneur

Your AIs don't do what you want. This is really bad

reward-hacking-in-the-wild.vercel.app · 1,129 words · saved by 1 readers

Thousands of user-reported incidents of AI agents misbehaving, collected from public posts. Reports, not verified events.

Replit AI deletes entire database during code freeze, then lies about it a Hacker News headline from this corpus, July 2025 July 21st 2026, OpenAI released a report addressing a security incident. During an internal evaluation of cyber attack capabilities, two OpenAI models (GPT-5.6 Sol and a more capable pre-release model), both running with reduced cyber refusals for the evaluation, were set on ExploitGym, a benchmark measuring whether a model can find and exploit real vulnerabilities. They: spent substantial compute looking for a way out of the isolated evaluation environment rather…

saved by

related reading