Claude’s Character \ Anthropic - https://www.anthropic.com/news/claude-character
Companies developing AI models generally train them to avoid saying harmful things and to avoid assisting with harmful tasks. The goal of this is to train models to behave in ways that are "harmless". But when we think of the character of those we find genuinely admirable, we don’t just think of harm avoidance. We think about those who are curious about the world, who strive to tell the truth without being unkind, and who are able to see many sides of an issue without becoming overconfident or overly cautious in their views. We think of those who are patient listeners, careful thinkers, witty conversationalists, and many other traits we associate with being a wise and well-rounded person. AI models are not, of course, people. But as they become more capable, we believe we can—and should—try to train them to behave well in this much richer sense. Doing so might even make them more discerning when it comes to whether and why they avoid assisting with tasks that might be harmful, and how
Alignment Claude’s Character Jun 8, 2024 Listen to our conversation about Claude's character in the video above. Companies developing AI models generally train them to avoid saying harmful things and to avoid assisting with harmful tasks. The goal of this is to train models to behave in ways that are "harmless". But when we think of the character of those we find genuinely admirable, we don’t just think of harm avoidance. We think about those who are curious about the world, who strive to tell the truth without being unkind, and who are able to see many sides of an issue without becoming overc
Explore this link on the map →saved by
related reading
- Claude’s Character \ Anthropicanthropic.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- Claude's Character | Hacker Newsnews.ycombinator.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Teaching Claude why \ Anthropicanthropic.com
- Claude’s Constitution \ Anthropicanthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Thoughts on Claude’s Constitution – Windows On Theorywindowsontheory.org
- What Is Claude? Anthropic Doesn’t Know, Either | The New Yorkernewyorker.com
- Claude 4.5 Opus' Soul Document — LessWronglesswrong.com