2211.06869
arxiv.org · 6,243 words · saved by 1 readers
N/A
Large Language Models Meet Harry Potter: A Bilingual Dataset for Aligning Dialogue Agents with Characters Nuo Chen†∗, Yan Wang‡§ , Haiyun Jiang‡ , Deng Cai‡ Yuhan Li† , Ziyang Chen‡ , Longyue Wang‡ and Jia Li†§ ‡ Tencent AI Lab †…
saved by
related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Opening the character training pipeline - by Nathan Lambertinterconnects.ai
- 2305.07759arxiv.org
- The persona selection model — LessWronglesswrong.com
- Claude’s Character \ Anthropicanthropic.com
- Unsupervised Elicitationalignment.anthropic.com
- PostTrainBenchposttrainbench.com