Emergent Dexterity via Diverse Resets and Large-Scale Reinforcement Learning
Reinforcement learning in massively parallel physics simulations has driven major progress in sim-to-real robot learning. However, current approaches remain brittle and task-specific, relying on extensive per-task engineering to design rewards, curricula, and demonstrations. Even with this engineering, typical reinforcement learning methods can often fail on long-horizon, contact-rich manipulation tasks and do not meaningfully scale with compute, as performance quickly saturates when training revisits the same narrow regions of state space. We introduce OmniReset, a simple and scalable framework that enables on-policy reinforcement learning to robustly solve a broad class of dexterous manipulation tasks using fixed algorithm hyperparameters, no curricula, minimal reward engineering, and no human demonstrations. Our key insight is that long-horizon exploration can be dramatically simplified by using simulator resets to systematically expose the RL algorithm to the diverse set of robot-o
Tyler Westenbroek Affiliation: University of Washington Octi Zhang Affiliation: NVIDIA Joshua Tran Affiliation: University of Washington Ignacio Dagnino Affiliation: University of Washington Eeshani Shilamkar Affiliation: University of Washington Numfor Mbiziwo-Tiapo Affiliation: University of Washington Simran Bagaria Affiliation: Microsoft Research*Equal contribution Xinlei Liu Affiliation: University of Washington Galen Mullins Affiliation: Microsoft Research*Equal contribution Andrey Kolobov Affiliation: Microsoft Research*Equal contribution Abhishek Gupta…
saved by
related reading
- Ch. 21 - Imitation Learningunderactuated.mit.edu
- State of Robot Learning, December 2025vedder.io
- Learning dexterity | OpenAIopenai.com
- Learning Sim-to-Real Humanoid Locomotion in 15 Minutesyounggyo.me
- Precise Manipulation with Efficient Online RLpi.website
- Building Worlds That Train Robotsworldlabs.ai
- Stop Simulating, Start Experiencingpaoloai.substack.com
- Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learningarxiv.org
- ALOHA Unleashed: A Simple Recipe for Robot Dexterityaloha-unleashed.github.io
- Behavioral cloning mysteryseohong.me
- Local Policies Enable Zero-shot Long-horizon Manipulationarxiv.org
- A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulationtoyotaresearchinstitute.github.io