Reward Is Not Enough - LessWrong
THREE CASE STUDIES 1. INCENTIVE LANDSCAPES THAT CAN’T FEASIBLY BE INDUCED BY A REWARD FUNCTION You’re a deity, tasked with designing a bird brain. You want the bird to get good at singing, as judged…
x Reward Is Not Enough — LessWrong Corrigibility Neuroscience Reinforcement learning Subagents AI Frontpage 137 Reward Is Not Enough by Steven Byrnes 16th Jun 2021 AI Alignment Forum 12 min read 19 137 Ω 44 Three case studies 1. Incentive landscapes that can’t feasibly be induced by a reward function You’re a deity, tasked with designing a bird brain. You want the bird to get good at singing, as judged by a black-box hardcoded song-assessing algorithm that you already built into the brain last week. The bird chooses actions based in part on within-lifetime reinforcement learning involving dopa
Explore this link on the map →saved by
related reading
- Reward is not the optimization target — LessWronglesswrong.com
- Reward Function Design: a starter pack — LessWronglesswrong.com
- Models Don't "Get Reward" — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- The Era of Experience Paper.pdfstorage.googleapis.com
- Reward is not the optimization target — AI Alignment Forumalignmentforum.org
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Reward Is Not the Optimization Targetturntrout.com
- “Behaviorist” RL reward functions lead to scheming — AI Alignment Forumalignmentforum.org
- In (highly contingent!) defense of interpretability-in-the-loop ML training — AI Alignment Forumalignmentforum.org
- AGI safetysjbyrnes.com