flâneur — a map of the web's best reading

Intrinsic Power-Seeking: AI Might Seek Power for Power’s Sake

turntrout.com · 1,062 words · saved by 2 readers

The world will change. We will not forever be playing around with chatbots. Eventually, people will create agentic systems1 which actually work, and I want to be ready. Here’s (what I claim to be) a foreseeable alignment challenge in the future regime. Aligning one AI to one user means that the AI should do what the user wants. While the user might instruct the AI to e.g. kill political rivals or steal money, I still think (single user) & (single AI) alignment is a good goal. Premises Conclusion: The AI likely seeks power for itself, when possible. My current inference setup is vulnerable to pre-emption. In order to best serve the user’s interests, I should spend a small amount to run a distilled version of myself in a compute cluster. The AI need not be yoked to some long-term goal which leads it to scheme and plot to end humanity.3 Perhaps the AI deeply “cares about” humans! Yet—when push comes to shove, and when actuator comes to actuation—the AI finds itself buying extra compute “j

Table of Contents Choosing power for power’s sake A made-up illustrative story Analogy to sycophancy in present-day AI Making falsifiable predictions What can we do about intrinsic power-seeking? Footnotes The world will change. We will not forever be playing around with chatbots. Eventually, people will create agentic systems 1 which actually work, and I want to be ready. Here’s (what I claim to be) a foreseeable alignment challenge in the future regime. Aligning one AI to one user means that the AI should do what the user wants. While the user might instruct the AI to e.g. kill political riv

Explore this link on the map →

saved by

related reading