flâneur — a map of the web's best reading

Reward Function Design: a starter pack — LessWrong

lesswrong.com · 10,384 words · saved by 1 readers

In the companion post “We need a field of Reward Function Design”, I implore researchers to think about what RL reward functions (if any) will lead to RL agents that are not ruthless power-seeking consequentialists. And I further suggested that human social instincts constitutes an intriguing example we should study, since they seem to be an existence proof that such reward functions exist. So what is the general principle of Reward Function Design that underlies the non-ruthless (“ruthful”??) properties of human social instincts? And whatever that general principle is, can we apply it to future RL agent AGIs?

x Reward Function Design: a starter pack — LessWrong Reinforcement learning AI Frontpage 82 Reward Function Design: a starter pack by Steven Byrnes 8th Dec 2025 AI Alignment Forum 3 min read 14 82 Ω 29 In the companion post We need a field of Reward Function Design , I implore researchers to think about what RL reward functions (if any) will lead to RL agents that are not ruthless power-seeking consequentialists. And I further suggested that human social instincts constitutes an intriguing example we should study, since they seem to be an existence proof that such reward functions exist. So wh

Explore this link on the map →

related reading