Training on Documents About Reward Hacking Induces Reward Hacking — LessWrong
This is a blog post reporting some preliminary work from the Anthropic Alignment Science team, which might be of interest to researchers working acti…
x Training on Documents About Reward Hacking Induces Reward Hacking — LessWrong Self Fulfilling/Refuting Prophecies AI Frontpage 2025 Top Fifty: 5 % 135 Training on Documents About Reward Hacking Induces Reward Hacking by evhub , Nathan Hu 21st Jan 2025 AI Alignment Forum 2 min read 15 135 Ω 61 This is a linkpost for https://alignment.anthropic.com/2025/reward-hacking-ooc/ This is a blog post reporting some preliminary work from the Anthropic Alignment Science team, which might be of interest to researchers working actively in this space. We'd ask you to treat these results like those of a col
related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Training on Documents about Reward Hacking Induces Reward Hackingalignment.anthropic.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Training a Misaligned Reward Seekeralignment.anthropic.com
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- Maybe we should pretrain on synthetic data about good-but-reward-hacking AIs — LessWronglesswrong.com
- Simulated Users & Sad LLMs1a3orn.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- [2607.18966] Measuring Reward-Seeking via Contrastive Belief Updatesarxiv.org