Deep Reinforcement Learning from Human Preferences
proceedings.neurips.cc · 4,998 words · saved by 1 readers
N/A
# link_1v5y0rmw15d.pdf ## Metadata - PDFFormatVersion=1.3 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Subject=Neural Information Processing Systems http://nips.cc/ - Custom.Publisher=Curran Associates, Inc. - Custom.Language=en-US - Custom.Created=2017 - Custom.EventType=Poster - Custom.Description-Abstract=For sophisticated reinforcement learning (RL) systems to interact usefully with real-world environments, we need to communicate complex goals to these systems. In this work, we explore goals defined in terms o
saved by
related reading
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Reinforcement learning from human feedback - Wikipediaen.wikipedia.org
- Learning through human feedback — Google DeepMinddeepmind.google
- Reward is not the optimization target — LessWronglesswrong.com
- RLHF Bookrlhfbook.com
- RLHF | John Lambertjohnwlambert.github.io
- [2309.00267] RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedbackarxiv.org
- Deep Reinforcement Learning: Pong from Pixelskarpathy.github.io
- Deep Reinforcement Learning Doesn't Work Yetalexirpan.com
- rlhfbook.com/book.pdfrlhfbook.com