Training on Documents about Reward Hacking Induces Reward Hacking
This is a blog post reporting some preliminary work from the Anthropic Alignment Science team, which might be of interest to researchers working actively in this space. We'd ask you to treat these results like those of a colleague sharing some thoughts or preliminary experiments at a lab meeting, rather than a mature paper. Specifically, we report a demonstration of a form of Out-of-Context Reasoning where training on documents which discuss (but don’t demonstrate) Claude’s tendency to reward hack can lead to an increase or decrease in reward hacking behavior. In this work, we investigate the extent to which pretraining datasets can influence the higher-level behaviors of large language models (LLMs). While pretraining shapes the factual knowledge and capabilities of LLMs (Petroni et al. 2019, Roberts et al. 2020, Lewkowycz et al. 2022, Allen-Zhu & Li, 2023), it is less well-understood whether it also affects their demonstrated preferences. We study whether training documents discussin
Training on Documents about Reward Hacking Induces Reward Hacking Alignment Science Blog Training on Documents About Reward Hacking Induces Reward Hacking Nathan Hu Benjamin Wright, Carson Denison, Samuel Marks, Johannes Treutlein, Jonathan Uesato Evan Hubinger This is a blog post reporting some preliminary work from the Anthropic Alignment Science team, which might be of interest to researchers working actively in this space. We'd ask you to treat these results like those of a colleague sharing some thoughts or preliminary experiments at a lab meeting, rather than a mature paper. Specifically
Explore this link on the map →related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Training on Documents About Reward Hacking Induces Reward Hacking — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Systematic Reward Hacking and Prime Sprintsprimeintellect.ai
- Training a Reward Hacker Despite Perfect Labels — LessWronglesswrong.com
- Natural emergent misalignment from reward hacking \ Anthropicanthropic.com
- Measuring Reward-Seeking by Instilling Contrastive Beliefsalignment.openai.com