Systematic Reward Hacking and Prime Sprints
We release tunable RL templates that demonstrate reward hacking at 1B scale and introduce Prime Sprints, an open-access program with sponsored runs for community research.
Systematic Reward Hacking and Prime Sprints Detecting and mitigating reward hacking is one of the key challenges faced when scaling RL, particularly in semi-verifiable domains. However, we lack systematic methods to understand when and why hacks emerge. Traditional wisdom describes reward hacking as a specification problem, where reward functions are simply too vague or not robust enough, and models inevitably learn to find exploits. While partially true, this offers little in the way of remediation other than “just make your rewards better”. From our experiences deploying RL across many domai
Explore this link on the map →saved by
related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- The Reward Hacking Benchmarkkunvarthaman.com
- Training on Documents about Reward Hacking Induces Reward Hackingalignment.anthropic.com
- Through the looking glass of benchmark hacking — Poolsidepoolside.ai
- Paper: Prompt Optimization Makes Misalignment Legible — LessWronglesswrong.com
- Training a Reward Hacker Despite Perfect Labels — LessWronglesswrong.com
- [2605.12474] Reward Hacking in Rubric-Based Reinforcement Learningarxiv.org