[1906.09624] On the Feasibility of Learning, Rather than Assuming, Human Biases for Reward Inference
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[1906.09624] On the Feasibility of Learning, Rather than Assuming, Human Biases for Reward Inference Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Machine Learning arXiv:1906.09624 (cs) [Submitted on 23 Jun 2019] Title: On the Feasibility of Learning, Rather than Assuming, Human Biases for Reward Inference Authors: Rohin Shah , Noah Gundotra , Pieter Abbeel , Anca D. Dragan View a PDF of the paper titled On the Feasibility of Learning, Rather than Assuming, Human Biases for Rewar
Explore this link on the map →related reading
- [1712.05812] Occam's razor is insufficient to infer the preferences of irrational agentsarxiv.org
- Reward is not the optimization target — LessWronglesswrong.com
- [2208.10687] The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Typesar5iv.labs.arxiv.org
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- [1712.05812] Occam's razor is insufficient to infer the preferences of irrational agentsarxiv.org
- Deep Reinforcement Learning from Human Preferencesproceedings.neurips.cc
- The Era of Experience Paper.pdfstorage.googleapis.com
- Models Don't "Get Reward" — LessWronglesswrong.com
- [AN #70]: Agents that help humans who are still learning about their own preferences — LessWronglesswrong.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Reward Is Not Enough — LessWronglesswrong.com