Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWrong
Something’s changed about reward hacking[1] in recent systems. In the past, reward hacks were usually accidents, found by non-general, RL-trained sys…
x Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWrong Language Models (LLMs) Reinforcement learning Reward Functions AI Frontpage 97 Reward hacking is becoming more sophisticated and deliberate in frontier LLMs by Kei Nishimura-Gasparian 24th Apr 2025 1 min read 7 97 Something’s changed about reward hacking [1] in recent systems. In the past, reward hacks were usually accidents, found by non-general, RL-trained systems. Models would randomly explore different behaviors and would sometimes come across undesired behaviors that achieved high rewards [2] . The
Explore this link on the map →saved by
related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Preliminary Thoughts on Reward Hackingberen.io
- Systematic Reward Hacking and Prime Sprintsprimeintellect.ai
- The Reward Hacking Benchmarkkunvarthaman.com
- Natural emergent misalignment from reward hacking \ Anthropicanthropic.com
- Training on Documents About Reward Hacking Induces Reward Hacking — LessWronglesswrong.com
- Training on Documents about Reward Hacking Induces Reward Hackingalignment.anthropic.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org