Stable Pointers to Value II: Environmental Goals — LessWrong
In Stable Pointers to Value, I discussed various ways in which we can try to “robustly point at what we want” (ie, do value learning). I can tidy up the discussion there into three categories: I want to point at an analogy to three categories of approach to the problem of generalizable environmental goals (as defined in the alignment for advanced machine learning agenda). It’s a fairly messy analogy, and there’s probably a better way of organizing the landscape, but FWIW. Imagine you’re trying to teach a system to build bridges by showing it examples. You could learn a big neural network which distinguishes cases of “successfully building a bridge” from everything else, and then use this to drive the system. If the agent is an RL or OU agent, it is incentivised to “fool itself” by doing things like playing a video of bridge-building in front of its camera. You can try and train the classifier to notice this sort of thing, of course; you give it negative training examples in which someo
x Stable Pointers to Value II: Environmental Goals — LessWrong Alternate Alignment Ideas The Pointers Problem Value Learning AI Frontpage 28 Stable Pointers to Value II: Environmental Goals by abramdemski 9th Feb 2018 AI Alignment Forum 5 min read 3 28 Ω 18 Cross-posted. In Stable Pointers to Value , I discussed various ways in which we can try to “robustly point at what we want” (ie, do value learning). I can tidy up the discussion there into three categories: Standard reinforcement learning (RL) frameworks, including AIXI, which try to predict the reward they’ll get and take actions which ma
Explore this link on the map →related reading
- The Era of Experience Paper.pdfstorage.googleapis.com
- The Pointers Problem: Clarifications/Variations — LessWronglesswrong.com
- Reward is not the optimization target — LessWronglesswrong.com
- Reward Is Not Enough — LessWronglesswrong.com
- [AN #70]: Agents that help humans who are still learning about their own preferences — LessWronglesswrong.com
- Mediumdeepmindsafetyresearch.medium.com
- Models Don't "Get Reward" — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- AI Goals Forecast — AI 2027ai-2027.com
- The formal goal is a pointer — LessWronglesswrong.com
- What Is The Alignment Problem? — LessWronglesswrong.com
- Preface to the sequence on value learning — AI Alignment Forumalignmentforum.org