Training language models to follow instructions with human feedback.pdf
proceedings.neurips.cc · 8,276 words · saved by 1 readers
N/A
Training language models to follow instructions with human feedback Long Ouyang∗ Jeff Wu∗ Xu Jiang∗ Diogo Almeida∗ Carroll L. Wainwright∗ Pamela Mishkin∗ Chong Zhang Sandhini Agarwal Katarina Slama Alex Ray John Schulman Jacob Hilton Fraser Kelton Luke Miller Maddie Simens Amanda Askell† Peter Welinder Paul Christiano∗† Jan Leike∗ Ryan Lowe∗…
saved by
related reading
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Checklists Are Better Than Reward Models For Aligning Language Modelsmachinelearning.apple.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- 2401.10020.pdfarxiv.org
- Stanford CRFMcrfm.stanford.edu
- [2005.14165] Language Models are Few-Shot Learnersarxiv.org
- 2308.03958arxiv.org
- Unsupervised Elicitationalignment.anthropic.com
- Scaling Instruction-Finetuned Language Modelsarxiv.org
- Unsupervised Elicitation of Language Modelsarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com