flâneur

Complex behavior from intrinsic motivation to occupy future action-state path space | Nature Communications

nature.com · 9,267 words · saved by 1 readers

Intelligent behavior of artificial agents and their design are usually considered as a reward maximization phenomenon, however, the reward function construction may be challenging. The authors introduce an alternative principle for agents’ behavior and design based on maximizing the occupancy of possible state and action paths.

Introduction Natural agents are endowed with a tendency to move, explore, and interact with their environment1,2. For instance, human newborns unintentionally move their body parts3, and 7–12-month-old infants spontaneously babble vocally4 and with their hands5. Exploration and curiosity are major drives for learning and discovery through information-seeking6,7,8. These behaviors seem to elude a simple explanation in terms of extrinsic reward maximization. However, intrinsic motivations, such as curiosity, push agents to visit new states by performing novel courses of action, which helps…

saved by

related reading