Are AIs more likely to pursue on-episode or beyond-episode reward?
blog.redwoodresearch.org · 2,316 words · saved by 1 readers
RL would encourage on-episode reward seeking, but beyond-episode reward seekers may learn to goal-guard.
Consider an AI that terminally pursues reward. How dangerous is this? It depends on how broadly-scoped a notion of reward the model pursues. It could be: on-episode reward-seeking: only maximizing reward on the current training episode — i.e., reward that reinforces their current action in RL. This is what people usually mean by “reward-seeker” (e.g. in Carlsmith or The behavioral selection model…). beyond-episode reward-seeking: maximizing reward for a larger-scoped notion of “self” (e.g., all models sharing the same weights). In this post, I’ll discuss which motivation is more likely.…
related reading
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Fitness-Seekers: Generalizing the Reward-Seeking Threat Model — LessWronglesswrong.com
- Reward is not the optimization target — LessWronglesswrong.com
- Fitness-Seekers: Generalizing the Reward-Seeking Threat Modelsubstack.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Training a Misaligned Reward Seekeralignment.anthropic.com
- The behavioral selection model for predicting AI motivationsblog.redwoodresearch.org
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org
- Models Don't "Get Reward" — LessWronglesswrong.com
- Reward Is Not the Optimization Targetturntrout.com
- Fail safe(r) at alignment by channeling reward-hacking into a "spillway" motivationblog.redwoodresearch.org