Open Problems of Reinforcement Learning
arxiv.org · 8,513 words · saved by 1 readers
N/A
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback Stephen Casper,∗ MIT CSAIL, scasper@mit.edu Xander Davies,∗ Harvard University Claudia Shi, Columbia University Thomas Krendl Gilbert, Cornell Tech Jérémy Scheurer, Apollo Research Javier Rando, ETH Zurich…
related reading
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- [2309.00267] RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedbackarxiv.org
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- rlhfbook.com/book.pdfrlhfbook.com
- Introduction | RLHF and Post-Training Book by Nathan Lambertrlhfbook.com
- Reinforcement learning from human feedback - Wikipediaen.wikipedia.org
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- Thoughts on the impact of RLHF research — LessWronglesswrong.com
- RLHF | John Lambertjohnwlambert.github.io
- [2402.14740] Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsarxiv.org
- How RLHF actually works - by Nathan Lambertinterconnects.ai
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io