Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWrong
Something’s changed about reward hacking[1] in recent systems. In the past, reward hacks were usually accidents, found by non-general, RL-trained sys…
x Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWrong Language Models (LLMs) Reinforcement learning Reward Functions AI Frontpage 97 Reward hacking is becoming more sophisticated and deliberate in frontier LLMs by Kei Nishimura-Gasparian 24th Apr 2025 1 min read 7 97 Something’s changed about reward hacking [1] in recent systems. In the past, reward hacks were usually accidents, found by non-general, RL-trained systems. Models would randomly explore different behaviors and would sometimes come across undesired behaviors that achieved high rewards [2] . The
saved by
related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Training a Misaligned Reward Seekeralignment.anthropic.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- Systematic Reward Hacking and Prime Sprintsprimeintellect.ai
- Preliminary Thoughts on Reward Hackingberen.io
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Simulated Users & Sad LLMs1a3orn.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- The Reward Hacking Benchmarkkunvarthaman.com