✳flâneur — a map of the web's best reading
The assistant axis: situating and stabilizing the character of large language models \ Anthropic
anthropic.com · 3,733 words · saved by 1 readers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Interpretability The assistant axis: situating and stabilizing the character of large language models Jan 19, 2026 Read the full paper Left: Character archetypes form a "persona space," with the Assistant at one extreme of the "Assistant Axis." Right: Capping drift along this axis prevents models (here, Llama 3.3 70B) from drifting into alternative personas and behaving in harmful ways. When you talk to a large language model, you can think of yourself as talking to a character . In the first stage of model training, pre-training, LLMs are asked to read vast amounts of text. Through this, they
Explore this link on the map →related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- [2601.10387] The Assistant Axis: Situating and Stabilizing the Default Persona of Language Modelsarxiv.org
- [2601.10387] The Assistant Axis: Situating and Stabilizing the Default Persona of Language Modelsarxiv.org
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- The persona selection model — LessWronglesswrong.com
- Claude’s Character \ Anthropicanthropic.com
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- Persona vectors: Monitoring and controlling character traits in language models \ Anthropicanthropic.com
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com