Reward is not the optimization target — LessWrong
TurnTrout discusses a common misconception in reinforcement learning: that reward is the optimization target of trained agents. He argues reward is b…
x Reward is not the optimization target — LessWrong Best of LessWrong 2022 Reinforcement learning Inner Alignment Reward Functions Wireheading Shard Theory Outer Alignment Deconfusion AI Frontpage 386 Reward is not the optimization target by TurnTrout 25th Jul 2022 AI Alignment Forum 12 min read 128 386 Ω 94 This insight was made possible by many conversations with Quintin Pope, where he challenged my implicit assumptions about alignment. I’m not sure who came up with this particular idea. In this essay, I call an agent a “reward optimizer” if it not only gets lots of reward, but if it reliabl
saved by
related reading
- Reward Is Not the Optimization Targetturntrout.com
- Reward is not the optimization target — AI Alignment Forumalignmentforum.org
- Models Don't "Get Reward" — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- Reward Is Not Enough — LessWronglesswrong.com
- How training-gamers might function (and win)blog.redwoodresearch.org
- [1912.01683] Optimal Policies Tend to Seek Powerarxiv.org
- Measuring Reward-Seeking by Instilling Contrastive Beliefsalignment.openai.com
- Evolution as Backstop for Reinforcement Learning · Gwern.netgwern.net
- The behavioral selection model for predicting AI motivationsblog.redwoodresearch.org
- [2607.18966] Measuring Reward-Seeking via Contrastive Belief Updatesarxiv.org