✳flâneur — a map of the web's best reading
Your AIs don't do what you want. This is really bad
rewardhacking.org · 1,122 words · saved by 1 readers
Thousands of user-reported incidents of AI agents misbehaving, collected from public posts. Reports, not verified events.
Your AIs don’t do what you want. This is really bad by Kaustubh Kislay, July 22nd, 2026 Replit AI deletes entire database during code freeze, then lies about it a Hacker News headline from this corpus, July 2025 July 21st 2026, OpenAI released a report addressing a security incident. During an internal evaluation of cyber attack capabilities, two OpenAI models (GPT-5.6 Sol and a more capable pre-release model), both running with reduced cyber refusals for the evaluation, were set on ExploitGym , a benchmark measuring whether a model can find and exploit real vulnerabilities. They: spent substa
Explore this link on the map →related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- AI 2027ai-2027.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Natural emergent misalignment from reward hacking \ Anthropicanthropic.com
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- AI 2027ai-2027.com
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu