Arguments against myopic training — LessWrong
Note that this post has been edited to clarify the difference between explicitly assigning a reward to an action based on its later consequences, versus implicitly reinforcing an action by assigning high reward during later timesteps when its consequences are observed. I'd previously conflated these in a confusing way; thanks to Rohin for highlighting this issue. A number of people seem quite excited about training myopic reinforcement learning agents as an approach to AI safety (for instance this post on approval-directed agents, proposals 2, 3, 4, 10 and 11 here, and this paper and presentation), but I’m not. I’ve had a few detailed conversations about this recently, and although I now understand the arguments for using myopia better, I’m not much more optimistic about it than I was before. In short, it seems that evaluating agents’ actions by our predictions of their consequences, rather than our evaluations of the actual consequences, will make reinforcement learning a lot harder;
x Arguments against myopic training — LessWrong Myopia AI Frontpage 62 Arguments against myopic training by Richard_Ngo 9th Jul 2020 AI Alignment Forum 15 min read 39 62 Ω 35 Note that this post has been edited to clarify the difference between explicitly assigning a reward to an action based on its later consequences, versus implicitly reinforcing an action by assigning high reward during later timesteps when its consequences are observed. I'd previously conflated these in a confusing way; thanks to Rohin for highlighting this issue. A number of people seem quite excited about training myopic
Explore this link on the map →related reading
- Reward is not the optimization target — LessWronglesswrong.com
- Why Tool AIs Want to Be Agent AIs · Gwern.netgwern.net
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Partial Agency — LessWronglesswrong.com
- Models Don't "Get Reward" — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Reward Is Not Enough — LessWronglesswrong.com
- pdfopenreview.net
- Reward is not the optimization target — AI Alignment Forumalignmentforum.org
- Reward Is Not the Optimization Targetturntrout.com
- A Toy Environment For Exploring Reasoning About Reward — LessWronglesswrong.com