Reward is not the optimization target — AI Alignment Forum
This insight was made possible by many conversations with Quintin Pope, where he challenged my implicit assumptions about alignment. I’m not sure who…
x Reward is not the optimization target — AI Alignment Forum Best of LessWrong 2022 Reinforcement learning Inner Alignment Reward Functions Wireheading Shard Theory Outer Alignment Deconfusion AI Frontpage 94 Reward is not the optimization target by TurnTrout 25th Jul 2022 12 min read 128 94 This insight was made possible by many conversations with Quintin Pope, where he challenged my implicit assumptions about alignment. I’m not sure who came up with this particular idea. In this essay, I call an agent a “reward optimizer” if it not only gets lots of reward, but if it reliably makes choices l
Explore this link on the map →related reading
- Reward is not the optimization target — LessWronglesswrong.com
- Reward Is Not the Optimization Targetturntrout.com
- Models Don't "Get Reward" — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward Is Not Enough — LessWronglesswrong.com
- Reward Function Design: a starter pack — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- A Toy Environment For Exploring Reasoning About Reward — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- Paper: Prompt Optimization Makes Misalignment Legible — LessWronglesswrong.com