Trajectory Improvement and Reward Learning from Comparative Language Feedback
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
Trajectory Improvement and Reward Learning from Comparative Language Feedback Zhaojing Yang University of Southern California Miru Jun University of Southern California Jeremy Tien University of California, Berkeley Stuart J. Russell University of California, Berkeley Anca Dragan University of California, Berkeley Erdem Bıyık University of Southern California Abstract Learning from human feedback has gained traction in fields like robotics and natural language processing in recent years. While prior works mostly rely on human feedback in the form of comparisons, language is a preferable modali
Explore this link on the map →related reading
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Deep Reinforcement Learning from Human Preferencesproceedings.neurips.cc
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- [2204.14146] Training Language Models with Language Feedbackarxiv.org
- [2301.02555] “No, to the Right” – Online Language Corrections for Robotic Manipulation via Shared Autonomyar5iv.labs.arxiv.org
- [2204.05186] Correcting Robot Plans with Natural Language Feedbackar5iv.labs.arxiv.org
- Reinforcement learning from human feedback - Wikipediaen.wikipedia.org
- rlhfbook.com/book.pdfrlhfbook.com
- Reward is not the optimization target — LessWronglesswrong.com