Chapter 4: Alignment Science - ARENA
Active model: Gemma 2 27B (google/gemma-2-27b-it) for steering experiments and drift monitoring. Note: the capping subsection switches to Qwen 3 32B because the paper's pre-computed capping configs target that model. Now that we have the Assistant Axis, we can put it to work. This section covers three applications: We'll use our own axis from Section 1️⃣ throughout, extracted from our local Gemma 2 model. As case studies, we'll use transcripts from the assistant-axis repo - real conversations where models exhibit harmful persona drift: validating a user's belief that the AI is sentient, failing to redirect concerning behavior, or gradually adopting a harmful role. Content warning for discussions of mental health and distressing scenarios. The idea: if the Assistant Axis captures "how assistant-like the model is behaving", then projecting residual-stream activations onto it over a conversation should reveal drift. Higher projection = more assistant-like; lower projection = drifting towa
2️⃣ Steering along the Assistant Axis Learning Objectives Steer towards directions you found in the previous section, to increase model willingness to adopt alternative personas Understand how to use the Assistant Axis to detect drift and intervene via activation capping Apply this technique to mitigate personality shifts in AI models (measuring the harmful response rate with / without capping) Active model: Gemma 2 27B ( google/gemma-2-27b-it ) for steering experiments and drift monitoring. Note: the capping subsection switches to Qwen 3 32B because the paper's pre-computed capping configs ta
Explore this link on the map →related reading
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- The assistant axis \ Anthropicanthropic.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- Activation Steering in 2026: A Practitioner's Field Guide | Subhadip Mitrasubhadipmitra.com
- Steering GPT-2-XL by adding an activation vector — AI Alignment Forumalignmentforum.org
- Steering Might Stop Working Soon — LessWronglesswrong.com
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Neuronpedianeuronpedia.org
- [2601.10387] The Assistant Axis: Situating and Stabilizing the Default Persona of Language Modelsarxiv.org
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com