flâneur — a map of the web's best reading

Training a Reward Hacker Despite Perfect Labels — LessWrong

lesswrong.com · 9,423 words · saved by 1 readers

Summary: Perfectly labeled outcomes in training can still boost reward hacking tendencies in generalization. This can hold even when the train/test sets are drawn from the exact same distribution. We induce this surprising effect via a form of context distillation, which we call re-contextualization: While we solely reinforce honest outcomes, the reasoning traces focus on hacking more than usual. We conclude that entraining hack-related reasoning boosts reward hacking. It's not enough to think about rewarding the right outcomes—we might also need to reinforce the right reasons. It's often thought that, if a model reward hacks on a task in deployment, then similar hacks were reinforced during training by a misspecified reward function.[1] In METR's report on reward hacking in frontier models, they posit the cause: "RL finds and reinforces strategies that receive high reward, and reward hacking is an effective strategy to get reward. In particular, the evaluation environments we’re usin

x Training a Reward Hacker Despite Perfect Labels — LessWrong Chain-of-Thought Alignment MATS Program AI Frontpage 2025 Top Fifty: 15 % 141 Training a Reward Hacker Despite Perfect Labels by ariana_azarbal , Victor Gillioz , TurnTrout 14th Aug 2025 AI Alignment Forum 5 min read 47 141 Ω 55 Summary: Perfectly labeled outcomes in training can still boost reward hacking tendencies in generalization. This can hold even when the train/test sets are drawn from the exact same distribution. We induce this surprising effect via a form of context distillation, which we call re-contextualization: Generat

Explore this link on the map →

related reading