Generally capable agents emerge from open-ended play - Google DeepMind
In recent years, artificial intelligence agents have succeeded in a range of complex game environments. For instance, AlphaZero beat world-champion programs in chess, shogi, and Go after starting out with knowing no more than the basic rules of how to play. Through reinforcement learning (RL), this single system learnt by playing round after round of games through a repetitive process of trial and error. But AlphaZero still trained separately on each game — unable to simply learn another game or task without repeating the RL process from scratch. The same is true for other successes of RL, such as Atari, Capture the Flag, StarCraft II, Dota 2, and Hide-and-Seek. DeepMind’s mission of solving intelligence to advance science and humanity led us to explore how we could overcome this limitation to create AI agents with more general and adaptive behaviour. Instead of learning one game at a time, these agents would be able to react to completely new conditions and play a whole universe of ga
Capture the Flag: the emergence of complex cooperative agents May 2019 Research Learn more
Explore this link on the map →related reading
- How DeepMind's Generally Capable Agents Were Trained — LessWronglesswrong.com
- DeepMind: Generally capable agents emerge from open-ended play — LessWronglesswrong.com
- The Era of Experience Paper.pdfstorage.googleapis.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Just Ask for Generalization | Eric Jangevjang.com
- [AN #159]: Building agents that know how to experiment, by training on procedurally generated games — AI Alignment Forumalignmentforum.org
- [AN #70]: Agents that help humans who are still learning about their own preferences — LessWronglesswrong.com
- Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver — LessWronglesswrong.com
- HANABI – np – ( ´ ▽ ` )ノnphard.io
- Learning Beyond Gradientstrinkle23897.github.io
- Voyager | An Open-Ended Embodied Agent with Large Language Modelsvoyager.minedojo.org
- Machine Studying | Jacob Xiaochen Lijacobxli.com