Chapter 4: Alignment Science - ARENA
Please send any problems / bugs on the #errata channel in the Slack group, and ask any questions on the dedicated channels for this chapter of material. If you want to change to dark mode, you can do this by clicking the three horizontal lines in the top-right, then navigating to Settings → Theme. Links to all other chapters: (0) Fundamentals, (1) Transformer Interpretability, (2) RL. Most exercises in this chapter have dealt with LLMs at quite a low level of abstraction: as mechanisms to perform certain tasks (e.g. indirect object identification, in-context antonym learning, or algorithmic tasks like predicting legal Othello moves). But if we want to study characteristics that have alignment relevance, we need a higher level of abstraction. LLMs often exhibit "personas" that can shift unexpectedly, sometimes dramatically (see Sydney, Grok's "MechaHitler" persona, or AI-induced psychosis). These personalities are clearly shaped through training and prompting, but exactly why remains a
[4.4] LLM Psychology & Persona Vectors Colab: exercises | solutions Please send any problems / bugs on the #errata channel in the Slack group , and ask any questions on the dedicated channels for this chapter of material. If you want to change to dark mode, you can do this by clicking the three horizontal lines in the top-right, then navigating to Settings → Theme. Links to all other chapters: (0) Fundamentals , (1) Transformer Interpretability , (2) RL . Introduction Most exercises in this chapter have dealt with LLMs at quite a low level of abstraction: as mechanisms to perform certain tasks
Explore this link on the map →related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- The persona selection model — LessWronglesswrong.com
- Persona vectors: Monitoring and controlling character traits in language models \ Anthropicanthropic.com
- [2507.21509] Persona Vectors: Monitoring and Controlling Character Traits in Language Modelsarxiv.org
- [2507.21509] Persona Vectors: Monitoring and Controlling Character Traits in Language Modelsarxiv.org
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- The assistant axis \ Anthropicanthropic.com
- [2601.10387] The Assistant Axis: Situating and Stabilizing the Default Persona of Language Modelsarxiv.org
- A Case for Model Persona Research — LessWronglesswrong.com
- [2601.10387] The Assistant Axis: Situating and Stabilizing the Default Persona of Language Modelsarxiv.org