Checklists Are Better Than Reward Models For Aligning Language Models - Apple Machine Learning Research
machinelearning.apple.com · 326 words · saved by 1 readers
Language models must be adapted to understand and follow user instructions. Reinforcement learning is widely used to facilitate this --…
AuthorsVijay Viswanathan†, Yanchao Sun, Shuang Ma‡**, Xiang Kong, Meng Cao, Graham Neubig†, Tongshuang Wu† Language models must be adapted to understand and follow user instructions. Reinforcement learning is widely used to facilitate this — typically using fixed criteria such as “helpfulness” and “harmfulness”. In our work, we instead propose using flexible, instruction-specific criteria as a means of broadening the impact that reinforcement learning can have in eliciting instruction following. We propose “Reinforcement Learning from Checklist Feedback” (RLCF). From instructions, we…
saved by
related reading
- Introduction | RLHF and Post-Training Book by Nathan Lambertrlhfbook.com
- user_interactions.pdfself-distillation.github.io
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Training language models to follow instructions with human feedback.pdfproceedings.neurips.cc
- 2401.10020.pdfarxiv.org
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Language Models can Solve Computer Tasksarxiv.org
- 2307.12950.pdfarxiv.org
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Stanford CRFMcrfm.stanford.edu