flâneur — a map of the web's best reading

Reward is not the optimization target — AI Alignment Forum

alignmentforum.org · 17,131 words · saved by 1 readers

This insight was made possible by many conversations with Quintin Pope, where he challenged my implicit assumptions about alignment. I’m not sure who…

x Reward is not the optimization target — AI Alignment Forum Best of LessWrong 2022 Reinforcement learning Inner Alignment Reward Functions Wireheading Shard Theory Outer Alignment Deconfusion AI Frontpage 94 Reward is not the optimization target by TurnTrout 25th Jul 2022 12 min read 128 94 This insight was made possible by many conversations with Quintin Pope, where he challenged my implicit assumptions about alignment. I’m not sure who came up with this particular idea. In this essay, I call an agent a “reward optimizer” if it not only gets lots of reward, but if it reliably makes choices l

Explore this link on the map →

related reading