[1907.10323] Fairness in Reinforcement Learning
Decision support systems (e.g., for ecological conservation) and autonomous systems (e.g., adaptive controllers in smart cities) start to be deployed in real applications. Although their operations often impact many users or stakeholders, no fairness consideration is generally taken into account in their design, which could lead to completely unfair outcomes for some users or stakeholders. To tackle this issue, we advocate for the use of social welfare functions that encode fairness and present this general novel problem in the context of (deep) reinforcement learning, although it could possibly be extended to other machine learning tasks.
Fairness in Reinforcement Learning Paul Weng Shanghai Jiao Tong University, Shanghai, China University of Michigan-Shanghai Jiao Tong University Joint Institute paul.weng@sjtu.edu.cn Abstract arXiv:1907.10323v1…
related reading
- [arxiv] an algorithmic framework for fairness elicitationarxiv.org
- Evolution as Backstop for Reinforcement Learning · Gwern.netgwern.net
- Reward is not the optimization target — LessWronglesswrong.com
- Papers · Nikhil Garggargnikhil.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- [1911.03020] A Human-in-the-loop Framework to Construct Context-aware Mathematical Notions of Outcome Fairnessarxiv.org
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Frontiers | On Consequentialism and Fairnessfrontiersin.org
- RLHF | John Lambertjohnwlambert.github.io
- Models Don't "Get Reward" — LessWronglesswrong.com
- Reinforcement Learning in Newcomblike Problemsproceedings.neurips.cc
- RLHF Bookrlhfbook.com