flâneur — a map of the web's best reading

Claude's Character | Hacker News

news.ycombinator.com · 11,378 words · saved by 1 readers

> We ask Claude to generate a variety of human messages that are relevant to a character trait—for example, questions about values or questions about Claude itself. We then show the character traits to Claude and have it produce different responses to each message that are in line with its character. Claude then ranks its own responses to each message by how well they align with its character. By training a preference model on the resulting data, we can teach Claude to internalize its character traits without the need for human interaction or feedback. You can't make a transformer-based language model "have" curiosity. Real curiosity is an innate trait in intelligent animals that promotes learning and exploration by increasing focus and stick-to-it-ness when the situation is interesting/unexplored (i.e. not well predicted - surprising). An LLM could be trained to fake curiosity in same way that ELIZA did ("Tell me more about your fear of prison showers"), but fundamentally it deals in

Claude's Character | Hacker News Hacker News new | past | comments | ask | show | jobs | submit login Claude's Character ( anthropic.com ) 263 points by simonw on June 9, 2024 | hide | past | favorite | 146 comments simonw on June 9, 2024 | next [–] > Claude 3 was the first model where we added "character training" to our alignment finetuning process: the part of training that occurs after initial model training, and the part that turns it from a predictive text model into an AI assistant. The goal of character training is to make Claude begin to have more nuanced, richer traits like curiosity

Explore this link on the map →

related reading