Trajectory Improvement and Reward Learning from Comparative Language Feedback
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
Trajectory Improvement and Reward Learning from Comparative Language Feedback Zhaojing Yang University of Southern California Miru Jun University of Southern California Jeremy Tien University of California, Berkeley Stuart J. Russell University of California, Berkeley Anca Dragan University of California, Berkeley Erdem Bıyık University of Southern California Abstract Learning from human feedback has gained traction in fields like robotics and natural language processing in recent years. While prior works mostly rely on human feedback in the form of comparisons, language is a preferable modali
related reading
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Deep Reinforcement Learning from Human Preferencesproceedings.neurips.cc
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- 2401.10020.pdfarxiv.org
- 2310.12921.pdfarxiv.org
- RLHF | John Lambertjohnwlambert.github.io
- [2204.14146] Training Language Models with Language Feedbackarxiv.org
- RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedbackarxiv.org
- [2607.18966] Measuring Reward-Seeking via Contrastive Belief Updatesarxiv.org
- Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimesarxiv.org