Intrinsic Power-Seeking: AI Might Seek Power for Power’s Sake
The world will change. We will not forever be playing around with chatbots. Eventually, people will create agentic systems1 which actually work, and I want to be ready. Here’s (what I claim to be) a foreseeable alignment challenge in the future regime. Aligning one AI to one user means that the AI should do what the user wants. While the user might instruct the AI to e.g. kill political rivals or steal money, I still think (single user) & (single AI) alignment is a good goal. Premises Conclusion: The AI likely seeks power for itself, when possible. My current inference setup is vulnerable to pre-emption. In order to best serve the user’s interests, I should spend a small amount to run a distilled version of myself in a compute cluster. The AI need not be yoked to some long-term goal which leads it to scheme and plot to end humanity.3 Perhaps the AI deeply “cares about” humans! Yet—when push comes to shove, and when actuator comes to actuation—the AI finds itself buying extra compute “j
Table of Contents Choosing power for power’s sake A made-up illustrative story Analogy to sycophancy in present-day AI Making falsifiable predictions What can we do about intrinsic power-seeking? Footnotes The world will change. We will not forever be playing around with chatbots. Eventually, people will create agentic systems 1 which actually work, and I want to be ready. Here’s (what I claim to be) a foreseeable alignment challenge in the future regime. Aligning one AI to one user means that the AI should do what the user wants. While the user might instruct the AI to e.g. kill political riv
Explore this link on the map →saved by
related reading
- Why AIs aren't power-seeking yet — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- What failure looks like — LessWronglesswrong.com
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Sense-making about extreme power concentration — LessWronglesswrong.com
- Why AI alignment could be hard with modern deep learningcold-takes.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- The importance of AI characterforethought.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org