flâneur — a map of the web's best reading

Chapter 4: Alignment Science - ARENA

learn.arena.education · 12,034 words · saved by 1 readers

Active model: Gemma 2 27B (google/gemma-2-27b-it) for steering experiments and drift monitoring. Note: the capping subsection switches to Qwen 3 32B because the paper's pre-computed capping configs target that model. Now that we have the Assistant Axis, we can put it to work. This section covers three applications: We'll use our own axis from Section 1️⃣ throughout, extracted from our local Gemma 2 model. As case studies, we'll use transcripts from the assistant-axis repo - real conversations where models exhibit harmful persona drift: validating a user's belief that the AI is sentient, failing to redirect concerning behavior, or gradually adopting a harmful role. Content warning for discussions of mental health and distressing scenarios. The idea: if the Assistant Axis captures "how assistant-like the model is behaving", then projecting residual-stream activations onto it over a conversation should reveal drift. Higher projection = more assistant-like; lower projection = drifting towa

2️⃣ Steering along the Assistant Axis Learning Objectives Steer towards directions you found in the previous section, to increase model willingness to adopt alternative personas Understand how to use the Assistant Axis to detect drift and intervene via activation capping Apply this technique to mitigate personality shifts in AI models (measuring the harmful response rate with / without capping) Active model: Gemma 2 27B ( google/gemma-2-27b-it ) for steering experiments and drift monitoring. Note: the capping subsection switches to Qwen 3 32B because the paper's pre-computed capping configs ta

Explore this link on the map →

related reading