2307.12950.pdf
arxiv.org · 8,212 words · saved by 1 readers
N/A
Published as a conference paper at ICLR 2024 RLCD: R EINFORCEMENT L EARNING FROM C ON - TRASTIVE D ISTILLATION FOR LM A LIGNMENT Kevin Yang1,2 Dan Klein2 Asli Celikyilmaz1 Nanyun Peng3 Yuandong Tian1 1 Meta AI, 2 UC Berkeley, 3 UCLA {yangk,klein}@berkeley.edu,{aslic,yuandong}@meta.com,violetpeng@cs.ucla.edu…
saved by
related reading
- 2310.12921.pdfarxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- RLHF | John Lambertjohnwlambert.github.io
- [2309.00267] RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedbackarxiv.org
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- rlhfbook.com/book.pdfrlhfbook.com
- [2402.14740] Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsarxiv.org
- Fine-tune Llama 2 with DPOhuggingface.co
- [2302.08582] Pretraining Language Models with Human Preferencesarxiv.org
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org