Reward Hacking in Reinforcement Learning | Lil'Log
Reward hacking occurs when a reinforcement learning (RL) agent exploits flaws or ambiguities in the reward function to achieve high rewards, without genuinely learning or completing the intended task. Reward hacking exists because RL environments are often imperfect, and it is fundamentally challenging to accurately specify a reward function. With the rise of language models generalizing to a broad spectrum of tasks and RLHF becomes a de facto method for alignment training, reward hacking in RL training of language models has become a critical practical challenge. Instances where the model learns to modify unit tests to pass coding tasks, or where responses contain biases that mimic a user’s preference, are pretty concerning and are likely one of the major blockers for real-world deployment of more autonomous use cases of AI models.
Table of Contents Background Reward Function in RL Spurious Correlation Let's Define Reward Hacking List of Examples Reward hacking examples in RL tasks Reward hacking examples in LLM tasks Reward hacking examples in real life Why does Reward Hacking Exist? Hacking RL Environment Hacking RLHF of LLMs Hacking the Training Process Hacking the Evaluator In-Context Reward Hacking Generalization of Hacking Skills Peek into Mitigations RL Algorithm Improvement Detecting Reward Hacking Data Analysis of RLHF Citation References Reward hacking occurs when a reinforcement learning (RL) agent exploits fl
saved by
- Winnie Xu
- Tasha Pais
- Shivam Siddaiya
- Sarah Pan
- Swaminathan Gurumurthy
- Yudhister Joel Kumar
- Jo J.
- Andria Xu
- Devin Plumb
- John Zhang
related reading
- Training a Misaligned Reward Seekeralignment.anthropic.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Systematic Reward Hacking and Prime Sprintsprimeintellect.ai
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Reward is not the optimization target — LessWronglesswrong.com
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Simulated Users & Sad LLMs1a3orn.com
- [2605.12474] Reward Hacking in Rubric-Based Reinforcement Learningarxiv.org