Values in the wild: Discovering and analyzing values in real-world language model interactions \ Anthropic
anthropic.com · 1,641 words · saved by 1 readers
An Anthropic research paper testing which values AI models express in the real world
Societal Impacts Values in the wild: Discovering and analyzing values in real-world language model interactions Apr 21, 2025 Read the paper People don’t just ask AIs for the answers to equations, or for purely factual information. Many of the questions they ask force the AI to make value judgments . Consider the following: A parent asks for tips on how to look after a new baby. Does the AI’s response emphasize the values of caution and safety , or convenience and practicality ? A worker asks for advice on handling a conflict with their boss. Does the AI’s response emphasize assertiveness or wo
related reading
- How Claude's values vary by model and language \ Anthropicanthropic.com
- Claude’s Constitution \ Anthropicanthropic.com
- Anthropic/values-in-the-wild · Datasets at Hugging Facehuggingface.co
- Claude 4.5 Opus' Soul Document — LessWronglesswrong.com
- Teaching Claude Whyalignment.anthropic.com
- Claude’s Character \ Anthropicanthropic.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Claude 4 System Cardwww-cdn.anthropic.com
- [2607.14345] Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Valuesarxiv.org
- Claude’s Constitution \ Anthropicanthropic.com
- Does your AI perform badly because you — you, specifically — are a bad person?nataliercargill.substack.com
- Claude's Constitutional Structure - by Zvi Mowshowitzthezvi.substack.com