AI Goals Forecast — AI 2027
AI 2027 predicts that superhuman AIs will not be aligned to the values and goals intended by their human developers. This supplement justifies that assumption by discussing the possibilities for what goals the AIs end up with.
AI Goals Forecast — AI 2027 AI Goals Forecast Daniel Kokotajlo | April 2025 Introduction What goals will AIs have? This document attempts to taxonomize and explain the many possibilities, explain why each possibility is plausible, and then finally attempt to quantify our uncertainty and explain where our bottom-line guesses are coming from. Ultimately we had to make a specific choice to depict in our scenario, but we hope this document will convey the range of uncertainty we have. We are keen to get feedback on these hypotheses and the arguments surrounding them. What important considerations
Explore this link on the map →related reading
- AI 2027ai-2027.com
- AI 2027ai-2027.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- why assume AGIs will optimize for fixed goals? — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Beliefs are Chosen to Serve Goals — LessWronglesswrong.com
- AI 2027ai-2027.com
- Commentary on AGI Safety from First Principles — AI Alignment Forumalignmentforum.org
- Why AIs aren't power-seeking yet — LessWronglesswrong.com
- Why AI alignment could be hard with modern deep learningcold-takes.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org