The Pando Problem: Rethinking AI Individuality — LessWrong
Epistemic status: This post aims at an ambitious target: improving intuitive understanding directly. The model for why this is worth trying is that I believe we are more bottlenecked by people having good intuitions guiding their research than, for example, by the ability of people to code and run evals. Quite a few ideas in AI safety implicitly use assumptions about individuality that ultimately derive from human experience. When we talk about AIs scheming, alignment faking or goal preservation, we imply there is something scheming or alignment faking or wanting to preserve its goals or escape the datacentre. If the system in question were human, it would be quite clear what that individual system is. When you read about Reinhold Messner reaching the summit of Everest, you would be curious about the climb, but you would not ask if it was his body there, or his mind, or his motivations to climb like the “spirit of mountaineering” or some combination thereof. In humans, all the answers
x The Pando Problem: Rethinking AI Individuality — LessWrong AI Psychology AI Risk Concrete Stories Personal Identity Tiling Agents AI Frontpage 2025 Top Fifty: 12 % 134 The Pando Problem: Rethinking AI Individuality by Jan_Kulveit 28th Mar 2025 AI Alignment Forum 16 min read 14 134 Ω 42 Epistemic status: This post aims at an ambitious target: improving intuitive understanding directly. The model for why this is worth trying is that I believe we are more bottlenecked by people having good intuitions guiding their research than, for example, by the ability of people to code and run evals. Quite
Explore this link on the map →related reading
- The Artificial Self — LessWronglesswrong.com
- The Artificial Selftheartificialself.ai
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- AI 2027ai-2027.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Paths Forward — The Artificial Selftheartificialself.ai
- Foundations of Identity — The Artificial Selftheartificialself.ai
- Do Not Tile the Lightcone with Your Confused Ontology — LessWronglesswrong.com
- Cyborgism — LessWronglesswrong.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com