Reward hacking behavior can generalize across tasks — AI Alignment Forum
TL;DR: We find that reward hacking generalization occurs in LLMs in a number of experimental settings and can emerge from reward optimization on certain datasets. This suggests that when models exploit flaws in supervision during training, they can sometimes generalize to exploit flaws in supervision in out-of-distribution environments. Machine learning models can display reward hacking behavior, where models score highly on imperfect reward signals by acting in ways not intended by their designers. Researchers have hypothesized that sufficiently capable models trained to get high reward on a diverse set of environments could become general reward hackers. General reward hackers would use their understanding of human and automated oversight in order to get high reward in a variety of novel environments, even when this requires exploiting gaps in our evaluations and acting in ways we don’t intend. It appears likely that model supervision will be imperfect and incentivize some degree of
x Reward hacking behavior can generalize across tasks — AI Alignment Forum Reward Functions MATS Program AI Frontpage 45 Reward hacking behavior can generalize across tasks by Kei Nishimura-Gasparian , Isaac Dunn , Henry Sleight , Miles Turpin , evhub , Carson Denison , Ethan Perez 28th May 2024 25 min read 5 45 TL;DR: We find that reward hacking generalization occurs in LLMs in a number of experimental settings and can emerge from reward optimization on certain datasets. This suggests that when models exploit flaws in supervision during training, they can sometimes generalize to exploit flaws
Explore this link on the map →saved by
related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Systematic Reward Hacking and Prime Sprintsprimeintellect.ai
- Training a Reward Hacker Despite Perfect Labels — LessWronglesswrong.com
- Training on Documents about Reward Hacking Induces Reward Hackingalignment.anthropic.com
- Training on Documents About Reward Hacking Induces Reward Hacking — LessWronglesswrong.com
- The Reward Hacking Benchmarkkunvarthaman.com
- Paper: Prompt Optimization Makes Misalignment Legible — LessWronglesswrong.com