Training on Documents About Reward Hacking Induces Reward Hacking — LessWrong
This is a blog post reporting some preliminary work from the Anthropic Alignment Science team, which might be of interest to researchers working acti…
x Training on Documents About Reward Hacking Induces Reward Hacking — LessWrong Self Fulfilling/Refuting Prophecies AI Frontpage 2025 Top Fifty: 5 % 135 Training on Documents About Reward Hacking Induces Reward Hacking by evhub , Nathan Hu 21st Jan 2025 AI Alignment Forum 2 min read 15 135 Ω 61 This is a linkpost for https://alignment.anthropic.com/2025/reward-hacking-ooc/ This is a blog post reporting some preliminary work from the Anthropic Alignment Science team, which might be of interest to researchers working actively in this space. We'd ask you to treat these results like those of a col
Explore this link on the map →related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Training on Documents about Reward Hacking Induces Reward Hackingalignment.anthropic.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- Maybe we should pretrain on synthetic data about good-but-reward-hacking AIs — LessWronglesswrong.com
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- Natural emergent misalignment from reward hacking \ Anthropicanthropic.com
- Training a Reward Hacker Despite Perfect Labels — LessWronglesswrong.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com