Reward Function Design: a starter pack — LessWrong
In the companion post “We need a field of Reward Function Design”, I implore researchers to think about what RL reward functions (if any) will lead to RL agents that are not ruthless power-seeking consequentialists. And I further suggested that human social instincts constitutes an intriguing example we should study, since they seem to be an existence proof that such reward functions exist. So what is the general principle of Reward Function Design that underlies the non-ruthless (“ruthful”??) properties of human social instincts? And whatever that general principle is, can we apply it to future RL agent AGIs?
x Reward Function Design: a starter pack — LessWrong Reinforcement learning AI Frontpage 82 Reward Function Design: a starter pack by Steven Byrnes 8th Dec 2025 AI Alignment Forum 3 min read 14 82 Ω 29 In the companion post We need a field of Reward Function Design , I implore researchers to think about what RL reward functions (if any) will lead to RL agents that are not ruthless power-seeking consequentialists. And I further suggested that human social instincts constitutes an intriguing example we should study, since they seem to be an existence proof that such reward functions exist. So wh
Explore this link on the map →related reading
- Reward Is Not Enough — LessWronglesswrong.com
- Reward is not the optimization target — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward is not the optimization target — AI Alignment Forumalignmentforum.org
- “Behaviorist” RL reward functions lead to scheming — AI Alignment Forumalignmentforum.org
- Why Tool AIs Want to Be Agent AIs · Gwern.netgwern.net
- Models Don't "Get Reward" — LessWronglesswrong.com
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Reward Is Not the Optimization Targetturntrout.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- RL & search is a terrifying way to build AGI (an FAQ) — LessWronglesswrong.com
- [1711.02827] Inverse Reward Designarxiv.org