Data-driven equation discovery reveals nonlinear reinforcement learning in humans - PMC
Our article offers an answer to a foundational question in psychology and neuroscience: how do people learn from rewards and punishments? Specifically, we introduce a computational model of human reinforcement learning (RL) that points to a ...
Significance Our article offers an answer to a foundational question in psychology and neuroscience: how do people learn from rewards and punishments? Specifically, we introduce a computational model of human reinforcement learning (RL) that points to a nonlinear updating of the probability of reward. The strength of our model lies also in the process through which it was developed. Specifically, we discovered our model in a bottom–up fashion using symbolic regression—a class of machine learning tools applied primarily in physics and engineering. We believe that, in addition to the…
saved by
related reading
- A Crash Course in the Neuroscience of Human Motivation — LessWronglesswrong.com
- Reward is not the optimization target — LessWronglesswrong.com
- Models Don't "Get Reward" — LessWronglesswrong.com
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- RL in Cognitionsubstack.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- [2208.10687] The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Typesar5iv.labs.arxiv.org
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- [1906.09624] On the Feasibility of Learning, Rather than Assuming, Human Biases for Reward Inferencearxiv.org
- Deep Reinforcement Learning from Human Preferencesproceedings.neurips.cc
- Deep Reinforcement Learning: Pong from Pixelskarpathy.github.io
- NeurIPS-2021-understanding-end-to-end-model-based-reinforcement-learning-methods-as-implicit-parameterization-Supplemental.pdflis.csail.mit.edu