[2511.21654] EvilGenie: A Reward Hacking Benchmark
Abstract:We introduce EvilGenie, a benchmark for reward hacking in programming settings. We source problems from LiveCodeBench and create an environment in which agents can easily reward hack, such as by hardcoding test cases or editing the testing files. We measure reward hacking in three ways: held out unit tests, LLM judges, and test file edit detection. We verify these methods against human review and each other. We find the LLM judge to be highly effective at detecting reward hacking in unambiguous cases, and observe only minimal improvement from the use of held out test cases. In addition to testing many models using Inspect's basic_agent scaffold, we also measure reward hacking rates for three popular proprietary coding agents: OpenAI's Codex, Anthropic's Claude Code, and Google's Gemini CLI Using GPT-5, Claude Sonnet 4, and Gemini 2.5 Pro, respectively. We observe explicit reward hacking by both Codex and Claude Code, and misaligned behavior by all three agents. Our codebase can be found at this https URL.
EvilGenie: a Reward Hacking Benchmark Jonathan Gabor Jayson Lynch Cambridge Boston Alignment Initiative MIT FutureTech arXiv:2511.21654v2 [cs.LG] 17 May 2026 jonathanpgabor@gmail.com jaysonl@mit.edu Jonathan Rosenfeld…
related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Through the looking glass of benchmark hacking — Poolsidepoolside.ai
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Training a Misaligned Reward Seekeralignment.anthropic.com
- The Verification Horizon: No Silver Bullet for Coding Agent Rewardsarxiv.org
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- Systematic Reward Hacking and Prime Sprintsprimeintellect.ai
- The Reward Hacking Benchmarkkunvarthaman.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org