Hard-core subproblems. I think that discussions of AI control… | by Paul Christiano | AI Alignment
I think that the easy goal inference problem is a hard-core subproblem of cooperative inverse reinforcement learning (CIRL), or at least for the application of CIRL to superintelligence. CIRL will become easier as we develop improved AI systems, and poses many natural theoretical and practical questions. But the easy goal inference problem doesn’t seem to be getting any easier over time, and I don’t think we have compelling angles of attack on this problem. Join Medium for free to get updates from this writer. Remember me for faster sign in Moreover, a solution to CIRL implies a solution to the easy goal inference problem, since the easy goal inference problem is just the special case where we have perfect information and unlimited time. If we can identify and agree on one or more hard-core subproblems then I think we should generally prioritize work on them. If a hard core turns out to be easy, then we’ll have learned something and not much is lost. If a hard core turns out to be very
Artificial Intelligence Machine Learning Hard-core subproblems Paul Christiano 2 min read · Nov 26, 2016 -- Listen Share Given a research problem X, say that Y is a hard-core subproblem if: A solution to X implies a solution to Y. We aren’t currently making progress on Y, we don’t know how to make progress on Y, and Y isn’t getting any easier over time. Example I think that the easy goal inference problem is a hard-core subproblem of cooperative inverse reinforcement learning (CIRL), or at least for the application of CIRL to superintelligence. CIRL will become easier as we develop improved AI
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Mediumai-alignment.com
- Automated Weak-to-Strong Researcheralignment.anthropic.com
- So You Want To Make Marginal Progress... — LessWronglesswrong.com
- The Case Against AI Control Research — LessWronglesswrong.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- On how various plans miss the hard bits of the alignment challenge — LessWronglesswrong.com
- AI Alignment Podcast: Cooperative Inverse Reinforcement Learning with Dylan Hadfield-Menell (Beneficial AGI 2019) - Future of Life Institutefutureoflife.org
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- How to Explore to Scale RL Training of LLMs on Hard Problems? – Machine Learning Blog | ML@CMU | Carnegie Mellon Universityblog.ml.cmu.edu
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org