✳flâneur — a map of the web's best reading
Benefits of Assistance over Reward Learning | OpenReview
openreview.net · 31 words · saved by 1 readers
Much recent work has focused on how an agent can learn what to do from human feedback, leading to two major paradigms. The first paradigm is reward learning, in which the agent learns a reward...
Verifying your browser | OpenReview Verifying your browser Complete the check below to continue to OpenReview Please complete the verification above. Have an OpenReview account? Sign in to skip this check.
Explore this link on the map →related reading
- The Era of Experience Paper.pdfstorage.googleapis.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward is not the optimization target — LessWronglesswrong.com
- Deep Reinforcement Learning from Human Preferencesproceedings.neurips.cc
- Learning through human feedback — Google DeepMinddeepmind.google
- [1906.09624] On the Feasibility of Learning, Rather than Assuming, Human Biases for Reward Inferencearxiv.org
- Models Don't "Get Reward" — LessWronglesswrong.com
- Reward Is Not Enough — LessWronglesswrong.com
- [2208.10687] The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Typesar5iv.labs.arxiv.org
- [AN #70]: Agents that help humans who are still learning about their own preferences — LessWronglesswrong.com
- Reinforcement learning - Wikipediaen.wikipedia.org
- Trajectory Improvement and Reward Learning from Comparative Language Feedbackarxiv.org