Different senses in which two AIs can be “the same” — LessWrong
Sometimes people talk about two AIs being “the same” or “different” AIs. We think the intuitive binary of “same vs. different” conflates several conc…
x Different senses in which two AIs can be “the same” — LessWrong AI World Modeling Frontpage 83 Different senses in which two AIs can be “the same” by Vivek Hebbar , Buck 24th Jun 2024 AI Alignment Forum 5 min read 3 83 Ω 40 Sometimes people talk about two AIs being “the same” or “different” AIs. We think the intuitive binary of “same vs. different” conflates several concepts which are often better to disambiguate. In this post, we spell out some of these distinctions. We don’t think anything here is particularly novel; we wrote this post because we think it’s probably mildly helpful for peop
Explore this link on the map →saved by
related reading
- The Artificial Selftheartificialself.ai
- Role-playing vs Self-modelling — LessWronglesswrong.com
- Insights on Crosscoder Model Diffingtransformer-circuits.pub
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- The Artificial Self — LessWronglesswrong.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Your Left Brain Doesn't Trade With Your Right — LessWronglesswrong.com
- [2603.11353] The Artificial Self: Characterising the landscape of AI identityarxiv.org
- The Pando Problem: Rethinking AI Individuality — LessWronglesswrong.com
- From personas to intentions: towards a science of motivations for AI models — LessWronglesswrong.com
- Models Don't "Get Reward" — LessWronglesswrong.com