✳flâneur — a map of the web's best reading
Reward Is Not the Optimization Target
turntrout.com · 4,034 words · saved by 1 readers
RL doesn't train reward optimizers. Reward chisels cognition into agents. Worry less about safe objectives, more about good cognition.
Table of Contents Reward probably won’t be a deep RL agent’s primary optimization target The siren-like suggestiveness of the word “reward” When is reward the optimization target of the agent? Anticipated questions Dropping the old hypothesis Implications Citation Similar posts Appendix: The field of RL thinks reward is the optimization target Footnotes In this essay, I call an agent a “reward optimizer” if it not only gets lots of reward, but if it reliably makes choices like “reward but no task completion” (e.g. receiving reward without eating pizza) over “task completion but no reward” (e.g
Explore this link on the map →related reading
- Reward is not the optimization target — LessWronglesswrong.com
- Reward is not the optimization target — AI Alignment Forumalignmentforum.org
- Models Don't "Get Reward" — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward Is Not Enough — LessWronglesswrong.com
- Why Tool AIs Want to Be Agent AIs · Gwern.netgwern.net
- Reward Function Design: a starter pack — LessWronglesswrong.com
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- A Toy Environment For Exploring Reasoning About Reward — LessWronglesswrong.com
- [1912.01683] Optimal Policies Tend to Seek Powerarxiv.org