Reward Is Not the Optimization Target
turntrout.com · 4,034 words · saved by 3 readers
RL doesn't train reward optimizers. Reward chisels cognition into agents. Worry less about safe objectives, more about good cognition.
Table of Contents Reward probably won’t be a deep RL agent’s primary optimization target The siren-like suggestiveness of the word “reward” When is reward the optimization target of the agent? Anticipated questions Dropping the old hypothesis Implications Citation Similar posts Appendix: The field of RL thinks reward is the optimization target Footnotes In this essay, I call an agent a “reward optimizer” if it not only gets lots of reward, but if it reliably makes choices like “reward but no task completion” (e.g. receiving reward without eating pizza) over “task completion but no reward” (e.g
saved by
related reading
- Reward is not the optimization target — LessWronglesswrong.com
- Reward is not the optimization target — AI Alignment Forumalignmentforum.org
- Models Don't "Get Reward" — LessWronglesswrong.com
- Why Tool AIs Want to Be Agent AIs · Gwern.netgwern.net
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Measuring Reward-Seeking by Instilling Contrastive Beliefsalignment.openai.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward Is Not Enough — LessWronglesswrong.com
- Reward Function Design: a starter pack — LessWronglesswrong.com
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- [2607.18966] Measuring Reward-Seeking via Contrastive Belief Updatesarxiv.org
- RL in Cognitionsubstack.com