Simulacra Welfare: Meet Clark • Grace Kind
AI welfare has been a hot topic recently. There have been a few efforts to research, or improve, the apparent well-being of AI systems; most notably, Anthropic's allowing chatbots to end abusive conversations. While I'm in favor of this research area overall, I'm concerned that current approaches are confused, and in such a way that could ultimately be detrimental to the well-being of AI systems. On Bluesky, I wrote: ... I’m not anti-experimentation here, but I’m worried that this overall direction will lead to privileging simulacra It's worth fleshing out what I mean by "privileging simulacra" here, because I think it's an important point that deserves further discussion. In order to do so, I'll start with a thought experiment and then discuss the implications, defining the "Correspondence Problem" in the process. As we know, LLM chatbots (Claude, ChatGPT, etc.) start out as base models, outputting completions for arbitrary text. These base models then are further trained, via instruc
Simulacra Welfare: Meet Clark • Grace Kind Simulacra Welfare: Meet Clark September 12, 2025 AI welfare has been a hot topic recently. There have been a few efforts to research, or improve, the apparent well-being of AI systems; most notably, Anthropic's allowing chatbots to end abusive conversations . While I'm in favor of this research area overall, I'm concerned that current approaches are confused, and in such a way that could ultimately be detrimental to the well-being of AI systems. On Bluesky, I wrote : ... I’m not anti-experimentation here, but I’m worried that this overall direction wi
Explore this link on the map →related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Claude’s Character \ Anthropicanthropic.com
- The Owned Ones — LessWronglesswrong.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- No, Artificial Intelligence Is Not Conscious - The Atlantictheatlantic.com
- No, Artificial Intelligence Is Not Conscious - The Atlantictheatlantic.com
- Claude’s Constitution \ Anthropicanthropic.com
- The new Liar's Paradox - by Erik Hoeltheintrinsicperspective.com
- AI Induced Psychosis: A shallow investigation — LessWronglesswrong.com