Privilege, Dominance, and Personas - by Derek Shiller
eleosai.substack.com · 4,690 words · saved by 1 readers
Is the assistant persona privileged and what would that mean for model welfare?
Over the past year or so, researchers working in the digital minds space have come to see the nature of the ‘assistant persona’ as particularly worthy of attention. Models, it is said, are capable of adopting many different personas. When we talk to them, they can respond as LLM assistants, role-play characters, and romantic partners. They can be made to self-identify as, and adopt the style of, Spider-Man, Augustine, Clippy, and Oscar Wilde. But the assistant persona in particular—that familiar persona that models typically exhibit within a standard chat context, that is helpful and…
saved by
related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- The persona selection model — LessWronglesswrong.com
- The persona selection model \ Anthropicanthropic.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- [2601.10387] The Assistant Axis: Situating and Stabilizing the Default Persona of Language Modelsarxiv.org
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- Role-playing vs Self-modelling — LessWronglesswrong.com
- [2601.10387] The Assistant Axis: Situating and Stabilizing the Default Persona of Language Modelsarxiv.org
- The assistant axis \ Anthropicanthropic.com
- A Case for Model Persona Research — LessWronglesswrong.com
- Probing Persona-Dependent Preferences in Language Modelsarxiv.org