Generally capable agents emerge from open-ended play - Google DeepMind
In recent years, artificial intelligence agents have succeeded in a range of complex game environments. For instance, AlphaZero beat world-champion programs in chess, shogi, and Go after starting out with knowing no more than the basic rules of how to play. Through reinforcement learning (RL), this single system learnt by playing round after round of games through a repetitive process of trial and error. But AlphaZero still trained separately on each game — unable to simply learn another game or task without repeating the RL process from scratch. The same is true for other successes of RL, such as Atari, Capture the Flag, StarCraft II, Dota 2, and Hide-and-Seek. DeepMind’s mission of solving intelligence to advance science and humanity led us to explore how we could overcome this limitation to create AI agents with more general and adaptive behaviour. Instead of learning one game at a time, these agents would be able to react to completely new conditions and play a whole universe of ga
July 27, 2021 Research Open-Ended Learning Team In recent years, artificial intelligence agents have succeeded in a range of complex game environments. For instance, AlphaZero beat world-champion programs in chess, shogi, and Go after starting out with knowing no more than the basic rules of how to play. Through reinforcement learning (RL), this single system learnt by playing round after round of games through a repetitive process of trial and error. But AlphaZero still trained separately on each game — unable to simply learn another game or task without repeating the RL process from…
related reading
- How DeepMind's Generally Capable Agents Were Trained — LessWronglesswrong.com
- The Era of Experience Paper.pdfstorage.googleapis.com
- DeepMind: Generally capable agents emerge from open-ended play — LessWronglesswrong.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Just Ask for Generalization | Eric Jangevjang.com
- [AN #159]: Building agents that know how to experiment, by training on procedurally generated games — AI Alignment Forumalignmentforum.org
- As Rocks May Think | Eric Jangevjang.com
- Can AI Learn From Experience? EBR-Bench Results | Epoch AI | Epoch AIepoch.ai
- [AN #70]: Agents that help humans who are still learning about their own preferences — LessWronglesswrong.com
- TheEraOfExperience.pdfincompleteideas.net
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io