Reward hacking behavior can generalize across tasks — AI Alignment Forum
TL;DR: We find that reward hacking generalization occurs in LLMs in a number of experimental settings and can emerge from reward optimization on certain datasets. This suggests that when models exploit flaws in supervision during training, they can sometimes generalize to exploit flaws in supervision in out-of-distribution environments. Machine learning models can display reward hacking behavior, where models score highly on imperfect reward signals by acting in ways not intended by their designers. Researchers have hypothesized that sufficiently capable models trained to get high reward on a diverse set of environments could become general reward hackers. General reward hackers would use their understanding of human and automated oversight in order to get high reward in a variety of novel environments, even when this requires exploiting gaps in our evaluations and acting in ways we don’t intend. It appears likely that model supervision will be imperfect and incentivize some degree of
x Reward hacking behavior can generalize across tasks — AI Alignment Forum Reward Functions MATS Program AI Frontpage 45 Reward hacking behavior can generalize across tasks by Kei Nishimura-Gasparian , Isaac Dunn , Henry Sleight , Miles Turpin , evhub , Carson Denison , Ethan Perez 28th May 2024 25 min read 5 45 TL;DR: We find that reward hacking generalization occurs in LLMs in a number of experimental settings and can emerge from reward optimization on certain datasets. This suggests that when models exploit flaws in supervision during training, they can sometimes generalize to exploit flaws
saved by
related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Training a Misaligned Reward Seekeralignment.anthropic.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Systematic Reward Hacking and Prime Sprintsprimeintellect.ai
- Simulated Users & Sad LLMs1a3orn.com
- Training a Reward Hacker Despite Perfect Labels — LessWronglesswrong.com
- [2606.12016] Generalization Hacking: Models Can Game Reinforcement Learning by Preventing Behavioral Generalizationarxiv.org