[2109.11978] Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning
Abstract:In this work, we present and study a training set-up that achieves fast policy generation for real-world robotic tasks by using massive parallelism on a single workstation GPU. We analyze and discuss the impact of different training algorithm components in the massively parallel regime on the final policy performance and training times. In addition, we present a novel game-inspired curriculum that is well suited for training with thousands of simulated robots in parallel. We evaluate the approach by training the quadrupedal robot ANYmal to walk on challenging terrain. The parallel approach allows training policies for flat terrain in under four minutes, and in twenty minutes for uneven terrain. This represents a speedup of multiple orders of magnitude compared to previous work. Finally, we transfer the policies to the real robot to validate the approach. We open-source our training code to help accelerate further research in the field of learned legged locomotion.
View PDF HTML (experimental) Abstract:In this work, we present and study a training set-up that achieves fast policy generation for real-world robotic tasks by using massive parallelism on a single workstation GPU. We analyze and discuss the impact of different training algorithm components in the massively parallel regime on the final policy performance and training times. In addition, we present a novel game-inspired curriculum that is well suited for training with thousands of simulated robots in parallel. We evaluate the approach by training the quadrupedal robot ANYmal to walk on…
saved by
related reading
- Learning Sim-to-Real Humanoid Locomotion in 15 Minutesyounggyo.me
- State of Robot Learning, December 2025vedder.io
- How Claude Performs on Robotics Tasks \ Anthropicanthropic.com
- Emergent Dexterity via Diverse Resets and Large-Scale Reinforcement Learningarxiv.org
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work
- Precise Manipulation with Efficient Online RLpi.website
- Learning Beyond Gradientstrinkle23897.github.io
- Ch. 21 - Imitation Learningunderactuated.mit.edu
- Keep the Tokens Flowing: Lessons from 16 Open-Source RL Librarieshuggingface.co
- Supervised Policy Learning for Real Robotssupervised-robot-learning.github.io
- [1602.01783] Asynchronous Methods for Deep Reinforcement Learningarxiv.org
- Teacher Algorithms for Deep RL Agents that Generalize in Procedurally Generated Environments – Developmental Systems, a Blog of the Flowers Labdevelopmentalsystems.org