Inference-Time Reward Hacking in Large Language Models
arxiv.org · 7,539 words · saved by 1 readers
N/A
Inference-Time Reward Hacking in Large Language Models Hadi Khalaf∗ Claudio Mayrink Verdun Alex Oesterling Harvard University Harvard University Harvard University Himabindu Lakkaraju Flavio du Pin Calmon∗ arXiv:2506.19248v2 [cs.LG] 4 Nov 2025…
related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Training a Misaligned Reward Seekeralignment.anthropic.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Systematic Reward Hacking and Prime Sprintsprimeintellect.ai
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- [2607.18966] Measuring Reward-Seeking via Contrastive Belief Updatesarxiv.org
- How hard is it to inoculate against misalignment generalization? — LessWronglesswrong.com