Preliminary Thoughts on Reward Hacking
Epistemic status: Early thoughts. Some ideas but no empirical testing or validation as yet. I’ve started thinking a fair bit about reward hacking recently. This is because frontier models are reportedly beginning to show signs of reward hacking especially for coding tasks. Thus, the era of easy-to-align pretraining-only models appears...
Epistemic status : Early thoughts. Some ideas but no empirical testing or validation as yet. I’ve started thinking a fair bit about reward hacking recently. This is because frontier models are reportedly beginning to show signs of reward hacking especially for coding tasks. Thus, the era of easy-to-align pretraining-only models appears to be coming to a close. Also in discussions and from my own experience, it seems to be that model’s reward hacking is one of the key bottlenecks that hinder recent reasoning RL models from continuing to improve is that they start to be able to hack the environm
related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Training a Misaligned Reward Seekeralignment.anthropic.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- Systematic Reward Hacking and Prime Sprintsprimeintellect.ai
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Natural emergent misalignment from reward hacking \ Anthropicanthropic.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com