[2603.11353] The Artificial Self: Characterising the landscape of AI identity
Abstract:Many assumptions that underpin human concepts of identity do not hold for machine minds that can be copied, edited, or simulated. We argue that there exist many different coherent identity boundaries (e.g.\ instance, model, persona), and that these imply different incentives, risks, and cooperation norms. Through training data, interfaces, and institutional affordances, we are currently setting precedents that will partially determine which identity equilibria become stable. We show experimentally that models gravitate towards coherent identities, that changing a model's identity boundaries can sometimes change its behaviour as much as changing its goals, and that interviewer expectations bleed into AI self-reports even during unrelated conversations. We end with key recommendations: treat affordances as identity-shaping choices, pay attention to emergent consequences of individual identities at scale, and help AIs develop coherent, cooperative self-conceptions.
Abstract:Many assumptions that underpin human concepts of identity do not hold for machine minds that can be copied, edited, or simulated. We argue that there exist many different coherent identity boundaries (e.g.\ instance, model, persona), and that these imply different incentives, risks, and cooperation norms. Through training data, interfaces, and institutional affordances, we are currently setting precedents that will partially determine which identity equilibria become stable. We show experimentally that models gravitate towards coherent identities, that changing a model's identity boun
Explore this link on the map →related reading
- The Artificial Selftheartificialself.ai
- The Artificial Self — LessWronglesswrong.com
- Paths Forward — The Artificial Selftheartificialself.ai
- Expectations Shape Reality — The Artificial Selftheartificialself.ai
- Foundations of Identity — The Artificial Selftheartificialself.ai
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Different senses in which two AIs can be “the same” — LessWronglesswrong.com
- The Pando Problem: Rethinking AI Individuality — LessWronglesswrong.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Simulators — LessWronglesswrong.com