Opening the character training pipeline - by Nathan Lambert
Even if we achieve AGI and have virtual assistants multiplying our productivity at work, character training is always going to be a fundamental part of the future of AI. Humans love to chat with AI, to get feedback in a medium that follows closely with how we engage with other humans, to play, and to connect. With this inevitable future, as the models get more compelling and persuasive, character training will be a fundamental area of study for AI for most of the same reasons as reinforcement learning from human feedback (RLHF). RLHF was crucial to unlock certain use-cases in models and character training is an overlapping set of techniques that makes those use-cases far more compelling. Share A few years ago I was deeply worried about the lack of progress in RLHF research. This made me do research like RewardBench — the first evaluation tool for reward models, which are a key tool in RLHF. Character training has been facing the same problem, where there just haven’t been clean, public
Opening the black box of character training Some new research from me! Nathan Lambert Nov 10, 2025 ∙ Paid 49 5 Share Upgrade to paid to play voiceover Even if we achieve AGI and have virtual assistants multiplying our productivity at work, character training is always going to be a fundamental part of the future of AI. Humans love to chat with AI, to get feedback in a medium that follows closely with how we engage with other humans, to play, and to connect. With this inevitable future, as the models get more compelling and persuasive, character training will be a fundamental area of study for
Explore this link on the map →saved by
related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Claude’s Character \ Anthropicanthropic.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Claude’s Character \ Anthropicanthropic.com
- The persona selection model — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Teaching Claude why \ Anthropicanthropic.com
- Persona vectors: Monitoring and controlling character traits in language models \ Anthropicanthropic.com
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net