Unitree’s founder disputes VLA consensus, backs video-trained models for robotics
When most in the industry think of Unitree Robotics, they see a company focused on building robot hardware. But at the World Robot Conference (WRC), founder Wang Xingxing offered a different narrative. During his keynote at WRC, Wang dedicated a large portion of his talk to large models, algorithms, and data. It was a shift that didn’t go unnoticed. His comments sparked debate, particularly his critique of the vision-language-action (VLA) framework driving many of today’s embodied robots. Wang didn’t mince words: he called the VLA architecture “relatively dumb.” The main problem, in his view, is data, or the lack thereof. VLA models require vast, high-quality datasets to function effectively in the real world. While the scarcity of such data is widely acknowledged, many companies have pursued brute-force methods: gathering real-world robot data, generating simulation data, or building specialized data collection infrastructure. Wang believes that emphasis is misplaced. “People are payi
Describing VLA models as “relatively dumb,” he outlined an alternative approach to embodied intelligence. When most in the industry think of Unitree Robotics, they see a company focused on building robot hardware. But at the World Robot Conference (WRC) , founder Wang Xingxing offered a different narrative. During his keynote at WRC, Wang dedicated a large portion of his talk to large models, algorithms, and data. It was a shift that didn’t go unnoticed. His comments sparked debate, particularly his critique of the vision-language-action (VLA) framework driving many of today’s embodied robots.
Explore this link on the map →saved by
related reading
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- The Embodied Internetben.bolte.cc
- Fully autonomous robots are much closer than you think – Sergey Levinedwarkesh.com
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- how we accidentally solved robotics by watching 1 million hours of YouTube – atharva's blogksagar.bearblog.dev
- Android Dreamsandroid-dreams.ai
- The Final Offshoringfinaloffshoring.com
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- The Model That Dreams the Worldmoe-capital.com
- A VLA with Open-World Generalizationpi.website
- The flavor of the bitter lesson for computer vision - Vincent Sitzmannvincentsitzmann.com
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com