Complex behavior from intrinsic motivation to occupy future action-state path space | Nature Communications
Intelligent behavior of artificial agents and their design are usually considered as a reward maximization phenomenon, however, the reward function construction may be challenging. The authors introduce an alternative principle for agents’ behavior and design based on maximizing the occupancy of possible state and action paths.
Introduction Natural agents are endowed with a tendency to move, explore, and interact with their environment1,2. For instance, human newborns unintentionally move their body parts3, and 7–12-month-old infants spontaneously babble vocally4 and with their hands5. Exploration and curiosity are major drives for learning and discovery through information-seeking6,7,8. These behaviors seem to elude a simple explanation in terms of extrinsic reward maximization. However, intrinsic motivations, such as curiosity, push agents to visit new states by performing novel courses of action, which helps…
saved by
related reading
- go-explore-nature.pdfadrien.ecoffet.com
- Reward is not the optimization target — LessWronglesswrong.com
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Is Curiosity All You Need? On the Utility of Emergent Behaviours from Curious Explorationarxiv.org
- Intrinsically-Motivated Humans and Agents in Open-World Exploration | Aly Lidayanalyd.github.io
- A Crash Course in the Neuroscience of Human Motivation — LessWronglesswrong.com
- [1705.05363] Curiosity-driven Exploration by Self-supervised Predictionar5iv.labs.arxiv.org
- [1912.01683] Optimal Policies Tend to Seek Powerarxiv.org
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- Towards a scale-free theory of intelligent agencymindthefuture.info
- pdfopenreview.net
- Reward Is Not the Optimization Targetturntrout.com