flâneur — a map of the web's best reading

Reward is not the optimization target — LessWrong

lesswrong.com · 22,058 words · saved by 6 readers

TurnTrout discusses a common misconception in reinforcement learning: that reward is the optimization target of trained agents. He argues reward is b…

x Reward is not the optimization target — LessWrong Best of LessWrong 2022 Reinforcement learning Inner Alignment Reward Functions Wireheading Shard Theory Outer Alignment Deconfusion AI Frontpage 386 Reward is not the optimization target by TurnTrout 25th Jul 2022 AI Alignment Forum 12 min read 128 386 Ω 94 This insight was made possible by many conversations with Quintin Pope, where he challenged my implicit assumptions about alignment. I’m not sure who came up with this particular idea. In this essay, I call an agent a “reward optimizer” if it not only gets lots of reward, but if it reliabl

Explore this link on the map →

saved by

related reading