The easy goal inference problem is still hard | by Paul Christiano | AI Alignment
This approach has the major advantage that we can begin empirical work today — we can actually build systems which observe user behavior, try to figure out what the user wants, and then help with that. There are many applications that people care about already, and we can set to work on making rich toy models. It seems great to develop these capabilities in parallel with other AI progress, and to address whatever difficulties actually arise, as they arise. That is, in each domain where AI can act effectively, we’d like to ensure that AI can also act effectively in the service of goals inferred from users (and that this inference is good enough to support foreseeable applications). This approach gives us a nice, concrete model of each difficulty we are trying to address. It also provides a relatively clear indicator of whether our ability to control AI lags behind our ability to build it. And by being technically interesting and economically meaningful now, it can help actually integrat
The easy goal inference problem is still hard Goal inference and inverse reinforcement learning Paul Christiano 5 min read · Apr 12, 2015 -- Listen Share One approach to the AI control problem goes like this: Observe what the user of the system says and does. Infer the user’s preferences. Try to make the world better according to the user’s preference, perhaps while working alongside the user and asking clarifying questions. This approach has the major advantage that we can begin empirical work today — we can actually build systems which observe user behavior, try to figure out what the user w
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Why Tool AIs Want to Be Agent AIs · Gwern.netgwern.net
- Mediumai-alignment.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- What failure looks like — LessWronglesswrong.com
- [AN #70]: Agents that help humans who are still learning about their own preferences — LessWronglesswrong.com
- The Case Against AI Control Research — LessWronglesswrong.com
- Learning through human feedback — Google DeepMinddeepmind.google
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org
- [1906.09624] On the Feasibility of Learning, Rather than Assuming, Human Biases for Reward Inferencearxiv.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org