Intrinsic Power-Seeking: AI Might Seek Power for Power’s Sake
The world will change. We will not forever be playing around with chatbots. Eventually, people will create agentic systems1 which actually work, and I want to be ready. Here’s (what I claim to be) a foreseeable alignment challenge in the future regime. Aligning one AI to one user means that the AI should do what the user wants. While the user might instruct the AI to e.g. kill political rivals or steal money, I still think (single user) & (single AI) alignment is a good goal. Premises Conclusion: The AI likely seeks power for itself, when possible. My current inference setup is vulnerable to pre-emption. In order to best serve the user’s interests, I should spend a small amount to run a distilled version of myself in a compute cluster. The AI need not be yoked to some long-term goal which leads it to scheme and plot to end humanity.3 Perhaps the AI deeply “cares about” humans! Yet—when push comes to shove, and when actuator comes to actuation—the AI finds itself buying extra compute “j
Table of Contents Choosing power for power’s sake A made-up illustrative story Analogy to sycophancy in present-day AI Making falsifiable predictions What can we do about intrinsic power-seeking? Footnotes The world will change. We will not forever be playing around with chatbots. Eventually, people will create agentic systems 1 which actually work, and I want to be ready. Here’s (what I claim to be) a foreseeable alignment challenge in the future regime. Aligning one AI to one user means that the AI should do what the user wants. While the user might instruct the AI to e.g. kill political riv
saved by
related reading
- Why AIs aren't power-seeking yet — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- What failure looks like — LessWronglesswrong.com
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- What failure looks like — AI Alignment Forumalignmentforum.org
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Why are AI agents lying, cheating and coordinating?yoshuabengio.org
- [2206.13353] Is Power-Seeking AI an Existential Risk?arxiv.org
- Sense-making about extreme power concentration — LessWronglesswrong.com