Fail safe(r) at alignment by channeling reward-hacking into a "spillway" motivation
blog.redwoodresearch.org · 4,686 words · saved by 2 readers
A controlled reward-seeking motivation could make AI safer and more useful
It’s plausible that flawed RL processes will select for misaligned AI motivations.1 Some misaligned motivations are much more dangerous than others. So, developers should plausibly aim to control which kind of misaligned motivations emerge in this case. In particular, we tentatively propose that developers should try to make the most likely generalization of reward hacking a bespoke bundle of benign reward-seeking traits, called a spillway motivation. We call this process spillway design. We think spillway design could have two major benefits: Spillway design might decrease the probability…
saved by
related reading
- From personas to intentions: towards a science of motivations for AI models — LessWronglesswrong.com
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Risk from fitness-seeking AIs: mechanisms and mitigationssubstack.com
- Your AIs don't do what you want. This is really badrewardhacking.org
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Training a Misaligned Reward Seekeralignment.anthropic.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Fitness-Seekers: Generalizing the Reward-Seeking Threat Model — LessWronglesswrong.com
- Natural emergent misalignment from reward hacking \ Anthropicanthropic.com