flâneur — a map of the web's best reading

The assistant axis: situating and stabilizing the character of large language models \ Anthropic

anthropic.com · 3,733 words · saved by 1 readers

Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.

Interpretability The assistant axis: situating and stabilizing the character of large language models Jan 19, 2026 Read the full paper Left: Character archetypes form a "persona space," with the Assistant at one extreme of the "Assistant Axis." Right: Capping drift along this axis prevents models (here, Llama 3.3 70B) from drifting into alternative personas and behaving in harmful ways. When you talk to a large language model, you can think of yourself as talking to a character . In the first stage of model training, pre-training, LLMs are asked to read vast amounts of text. Through this, they

Explore this link on the map →

related reading