flâneur — a map of the web's best reading

Reward hacking behavior can generalize across tasks — AI Alignment Forum

alignmentforum.org · 10,034 words · saved by 1 readers

TL;DR: We find that reward hacking generalization occurs in LLMs in a number of experimental settings and can emerge from reward optimization on certain datasets. This suggests that when models exploit flaws in supervision during training, they can sometimes generalize to exploit flaws in supervision in out-of-distribution environments. Machine learning models can display reward hacking behavior, where models score highly on imperfect reward signals by acting in ways not intended by their designers. Researchers have hypothesized that sufficiently capable models trained to get high reward on a diverse set of environments could become general reward hackers. General reward hackers would use their understanding of human and automated oversight in order to get high reward in a variety of novel environments, even when this requires exploiting gaps in our evaluations and acting in ways we don’t intend. It appears likely that model supervision will be imperfect and incentivize some degree of

x Reward hacking behavior can generalize across tasks — AI Alignment Forum Reward Functions MATS Program AI Frontpage 45 Reward hacking behavior can generalize across tasks by Kei Nishimura-Gasparian , Isaac Dunn , Henry Sleight , Miles Turpin , evhub , Carson Denison , Ethan Perez 28th May 2024 25 min read 5 45 TL;DR: We find that reward hacking generalization occurs in LLMs in a number of experimental settings and can emerge from reward optimization on certain datasets. This suggests that when models exploit flaws in supervision during training, they can sometimes generalize to exploit flaws

Explore this link on the map →

saved by

related reading