DoorMan - NVIDIA GEAR Lab
We use a procedural generation pipeline that randomizes physical and visual properties of articulated objects: mass, handle type, hinge damping, stiffness, texture, background, etc. We apply GRPO fine-tuning to bootstrap the student on top of classical teacher-student distillation. We find this stage to be uniquely useful for challenging loco-manipulation tasks due to partial observability. The reward is mostly binary on the task success criteria. The result: 20-30% success rate improvement. Our policy demonstrates robust generalization to diverse real-world scenarios, successfully manipulating various door types under different environmental conditions. DoorMan on average completes the door opening task by up to 7.15 seconds faster than human teleoperators, who struggle with skillful loco-manipulation with articulated objects. While our policy demonstrates strong performance across diverse scenarios, we observe failure modes that highlight areas for future improvement. Common failure
DoorMan - NVIDIA GEAR Lab Loading Experience... DoorMan NVIDIA GEAR Team Home Paper Details arXiv Code Infinite Visual Randomizations Powered by IsaacLab RGB-Only Sim-to-Real Generalizable Policy Transfer Scroll to explore DoorMan Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy Transfer Haoru Xue 1,2,* , Tairan He 1,3,* , Zi Wang 1,* Qingwei Ben 1,4 , Wenli Xiao 1,3 , Zhengyi Luo 1 , Xingye Da 1 , Fernando Castañeda 1 , Guanya Shi 3 , Shankar Sastry 2 , Linxi "Jim" Fan 1,† , Yuke Zhu 1,† 1 NVIDIA, 2 UC Berkeley, 3 CMU, 4 CUHK * Equal Contribution, † Project Leads Mass Scale Si
Explore this link on the map →related reading
- State of Robot Learning, December 2025vedder.io
- Learning dexterity | OpenAIopenai.com
- A VLA with Open-World Generalizationpi.website
- [2509.22407] EMMA: Generalizing Real-World Robot Manipulation via Generative Visual Transferarxiv.org
- Precise Manipulation with Efficient Online RLpi.website
- e5b5c402bb7bd5e60bede6961d6fe39e-Paper-Conference.pdfproceedings.iclr.cc
- Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AIrhoda.ai
- FACTR: Force-Attending Curriculum Training for Contact-Rich Policy Learningjasonjzliu.com
- RLDG: Robotic Generalist Policy Distillation via Reinforcement Learningarxiv.org
- A Steerable Model with Emergent Capabilitiespi.website
- Ch. 21 - Imitation Learningunderactuated.mit.edu
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robotsarxiv.org