Training a Reward Hacker Despite Perfect Labels — LessWrong
Summary: Perfectly labeled outcomes in training can still boost reward hacking tendencies in generalization. This can hold even when the train/test sets are drawn from the exact same distribution. We induce this surprising effect via a form of context distillation, which we call re-contextualization: While we solely reinforce honest outcomes, the reasoning traces focus on hacking more than usual. We conclude that entraining hack-related reasoning boosts reward hacking. It's not enough to think about rewarding the right outcomes—we might also need to reinforce the right reasons. It's often thought that, if a model reward hacks on a task in deployment, then similar hacks were reinforced during training by a misspecified reward function.[1] In METR's report on reward hacking in frontier models, they posit the cause: "RL finds and reinforces strategies that receive high reward, and reward hacking is an effective strategy to get reward. In particular, the evaluation environments we’re usin
x Training a Reward Hacker Despite Perfect Labels — LessWrong Chain-of-Thought Alignment MATS Program AI Frontpage 2025 Top Fifty: 15 % 141 Training a Reward Hacker Despite Perfect Labels by ariana_azarbal , Victor Gillioz , TurnTrout 14th Aug 2025 AI Alignment Forum 5 min read 47 141 Ω 55 Summary: Perfectly labeled outcomes in training can still boost reward hacking tendencies in generalization. This can hold even when the train/test sets are drawn from the exact same distribution. We induce this surprising effect via a form of context distillation, which we call re-contextualization: Generat
Explore this link on the map →related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Training on Documents About Reward Hacking Induces Reward Hacking — LessWronglesswrong.com
- Systematic Reward Hacking and Prime Sprintsprimeintellect.ai
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Training on Documents about Reward Hacking Induces Reward Hackingalignment.anthropic.com
- Reward is not the optimization target — LessWronglesswrong.com
- Recontextualization Mitigates Specification Gaming Without Modifying the Specification — LessWronglesswrong.com