[2208.10687] The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types
When inferring reward functions from human behavior (be it demonstrations, comparisons, physical corrections, or e-stops), it has proven useful to model the human as making noisy-rational choices, with a “rationality c…
The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types Gaurav R. Ghosal 1 \equalcontrib , Matthew Zurek 2 \equalcontrib , Daniel S. Brown 3 , Anca D. Dragan 1 Abstract When inferring reward functions from human behavior (be it demonstrations, comparisons, physical corrections, or e-stops), it has proven useful to model the human as making noisy-rational choices, with a “rationality coefficient” capturing how much noise or entropy we expect to see in the human behavior. Prior work typically sets the rationality level to a constant value, regardless of th
Explore this link on the map →related reading
- A Crash Course in the Neuroscience of Human Motivation — LessWronglesswrong.com
- [1906.09624] On the Feasibility of Learning, Rather than Assuming, Human Biases for Reward Inferencearxiv.org
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward is not the optimization target — LessWronglesswrong.com
- Models Don't "Get Reward" — LessWronglesswrong.com
- Deep Reinforcement Learning from Human Preferencesproceedings.neurips.cc
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- [1712.05812] Occam's razor is insufficient to infer the preferences of irrational agentsarxiv.org
- [1712.05812] Occam's razor is insufficient to infer the preferences of irrational agentsarxiv.org
- pdfopenreview.net
- Feedbackloop-first Rationality — LessWronglesswrong.com