Claude's Character | Hacker News
> We ask Claude to generate a variety of human messages that are relevant to a character trait—for example, questions about values or questions about Claude itself. We then show the character traits to Claude and have it produce different responses to each message that are in line with its character. Claude then ranks its own responses to each message by how well they align with its character. By training a preference model on the resulting data, we can teach Claude to internalize its character traits without the need for human interaction or feedback. You can't make a transformer-based language model "have" curiosity. Real curiosity is an innate trait in intelligent animals that promotes learning and exploration by increasing focus and stick-to-it-ness when the situation is interesting/unexplored (i.e. not well predicted - surprising). An LLM could be trained to fake curiosity in same way that ELIZA did ("Tell me more about your fear of prison showers"), but fundamentally it deals in
Claude's Character | Hacker News Hacker News new | past | comments | ask | show | jobs | submit login Claude's Character ( anthropic.com ) 263 points by simonw on June 9, 2024 | hide | past | favorite | 146 comments simonw on June 9, 2024 | next [–] > Claude 3 was the first model where we added "character training" to our alignment finetuning process: the part of training that occurs after initial model training, and the part that turns it from a predictive text model into an AI assistant. The goal of character training is to make Claude begin to have more nuanced, richer traits like curiosity
Explore this link on the map →related reading
- Claude’s Character \ Anthropicanthropic.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Claude’s Character \ Anthropicanthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- Claude’s Constitution \ Anthropicanthropic.com
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- What Is Claude? Anthropic Doesn’t Know, Either | The New Yorkernewyorker.com
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- The persona selection model — LessWronglesswrong.com