Reward is not the optimization target — LessWrong
TurnTrout discusses a common misconception in reinforcement learning: that reward is the optimization target of trained agents. He argues reward is b…
x Reward is not the optimization target — LessWrong Best of LessWrong 2022 Reinforcement learning Inner Alignment Reward Functions Wireheading Shard Theory Outer Alignment Deconfusion AI Frontpage 386 Reward is not the optimization target by TurnTrout 25th Jul 2022 AI Alignment Forum 12 min read 128 386 Ω 94 This insight was made possible by many conversations with Quintin Pope, where he challenged my implicit assumptions about alignment. I’m not sure who came up with this particular idea. In this essay, I call an agent a “reward optimizer” if it not only gets lots of reward, but if it reliabl
Explore this link on the map →saved by
related reading
- Reward is not the optimization target — AI Alignment Forumalignmentforum.org
- Reward Is Not the Optimization Targetturntrout.com
- Models Don't "Get Reward" — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- [1912.01683] Optimal Policies Tend to Seek Powerarxiv.org
- Reward Is Not Enough — LessWronglesswrong.com
- Fitness-Seekers: Generalizing the Reward-Seeking Threat Model — LessWronglesswrong.com
- Optimality is the tiger, and agents are its teeth — LessWronglesswrong.com
- Evolution as Backstop for Reinforcement Learning · Gwern.netgwern.net
- Reward Function Design: a starter pack — LessWronglesswrong.com
- A Crash Course in the Neuroscience of Human Motivation — LessWronglesswrong.com