flâneur

Checklists Are Better Than Reward Models For Aligning Language Models - Apple Machine Learning Research

machinelearning.apple.com · 326 words · saved by 1 readers

Language models must be adapted to understand and follow user instructions. Reinforcement learning is widely used to facilitate this --…

AuthorsVijay Viswanathan†, Yanchao Sun, Shuang Ma‡**, Xiang Kong, Meng Cao, Graham Neubig†, Tongshuang Wu† Language models must be adapted to understand and follow user instructions. Reinforcement learning is widely used to facilitate this — typically using fixed criteria such as “helpfulness” and “harmfulness”. In our work, we instead propose using flexible, instruction-specific criteria as a means of broadening the impact that reinforcement learning can have in eliciting instruction following. We propose “Reinforcement Learning from Checklist Feedback” (RLCF). From instructions, we…

saved by

related reading