Reward Is Not Enough - LessWrong
THREE CASE STUDIES 1. INCENTIVE LANDSCAPES THAT CAN’T FEASIBLY BE INDUCED BY A REWARD FUNCTION You’re a deity, tasked with designing a bird brain. You want the bird to get good at singing, as judged…
x Reward Is Not Enough — LessWrong Corrigibility Neuroscience Reinforcement learning Subagents AI Frontpage 137 Reward Is Not Enough by Steven Byrnes 16th Jun 2021 AI Alignment Forum 12 min read 19 137 Ω 44 Three case studies 1. Incentive landscapes that can’t feasibly be induced by a reward function You’re a deity, tasked with designing a bird brain. You want the bird to get good at singing, as judged by a black-box hardcoded song-assessing algorithm that you already built into the brain last week. The bird chooses actions based in part on within-lifetime reinforcement learning involving dopa
saved by
related reading
- Reward is not the optimization target — LessWronglesswrong.com
- Reward Function Design: a starter pack — LessWronglesswrong.com
- Models Don't "Get Reward" — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Reward Is Not the Optimization Targetturntrout.com
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Reward is not the optimization target — AI Alignment Forumalignmentforum.org
- Fail safe(r) at alignment by channeling reward-hacking into a "spillway" motivationblog.redwoodresearch.org
- In (highly contingent!) defense of interpretability-in-the-loop ML training — AI Alignment Forumalignmentforum.org
- [1711.02827] Inverse Reward Designarxiv.org
- Deep Deceptiveness — LessWronglesswrong.com