✳flâneur — a map of the web's best reading
Emotion concepts and their function in a large language model \ Anthropic
anthropic.com · 2,958 words · saved by 8 readers
All modern language models sometimes act like they have emotions. What’s behind these behaviors? Our interpretability team investigates.
Interpretability Emotion concepts and their function in a large language model Apr 2, 2026 Read the paper All modern language models sometimes act like they have emotions. They may say they’re happy to help you, or sorry when they make a mistake. Sometimes they even appear to become frustrated or anxious when struggling with tasks. What’s behind these behaviors? The way modern AI models are trained pushes them to act like a character with human-like characteristics. In addition, these models are known to develop rich and generalizable internal representations of abstract concepts underlying th
Explore this link on the map →saved by
related reading
- Emotion Concepts and their Function in a Large Language Modeltransformer-circuits.pub
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- [2606.26987] Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMsarxiv.org
- Transformer Circuits Threadtransformer-circuits.pub
- A global workspace in language models \ Anthropicanthropic.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Mapping the mind of a large language model \ Anthropicanthropic.com
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org