What can AI Learn from Human Exploration? Intrinsically-Motivated Humans and Agents in Open-World Exploration | Aly Lidayan
What drives exploration? Understanding intrinsic motivation is a long-standing question in both cognitive science and artificial intelligence (AI); numerous exploration objectives have been proposed and tested in human experiments and used to train reinforcement learning (RL) agents. However, experiments in the former are often in simplistic environments that do not capture the complexity of real world exploration. On the other hand, experiments in the latter use more complex environments, yet the trained RL agents fail to come close to human exploration efficiency. To study this gap, we propose a framework for directly comparing human and agent exploration in an open-ended environment, Crafter. We study how well commonly-proposed information theoretic intrinsic objectives relate to actual human and agent behaviors, finding that human and intrinsically-motivated RL agent exploration success consistently show positive correlation with Entropy and Empowerment. However, only human explora
Intrinsically-Motivated Humans and Agents in Open-World Exploration | Aly Lidayan Search Intrinsically-Motivated Humans and Agents in Open-World Exploration Yuqing Du* , Eliza Kosoy* , Aly Lidayan* , Maria Rufova , Pieter Abbeel , Alison Gopnik November, 2023 PDF Abstract What drives exploration? Understanding intrinsic motivation is a long-standing challenge in both cognitive science and artificial intelligence; numerous objectives have been proposed and used to train agents, yet there remains a gap between human and agent exploration. We directly compare adults, children, and AI agents in a
Explore this link on the map →related reading
- What can AI Learn from Human Exploration? Intrinsically-Motivated Humans and Agents in Open-World Exploration – Center for Human-Compatible Artificial Intelligencehumancompatible.ai
- The Era of Experience Paper.pdfstorage.googleapis.com
- Why Tool AIs Want to Be Agent AIs · Gwern.netgwern.net
- AI 2027ai-2027.com
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- [1705.05363] Curiosity-driven Exploration by Self-supervised Predictionar5iv.labs.arxiv.org
- Exploration Strategies in Deep Reinforcement Learning | Lil'Loglilianweng.github.io
- Automated Weak-to-Strong Researcheralignment.anthropic.com
- Reward is not the optimization target — LessWronglesswrong.com
- Voyager | An Open-Ended Embodied Agent with Large Language Modelsvoyager.minedojo.org
- An Apple-Picking Model of AI R&D | Tom Cunningham – Tom Cunninghamtecunningham.github.io
- [1912.01683] Optimal Policies Tend to Seek Powerarxiv.org