Unitree’s founder disputes VLA consensus, backs video-trained models for robotics
When most in the industry think of Unitree Robotics, they see a company focused on building robot hardware. But at the World Robot Conference (WRC), founder Wang Xingxing offered a different narrative. During his keynote at WRC, Wang dedicated a large portion of his talk to large models, algorithms, and data. It was a shift that didn’t go unnoticed. His comments sparked debate, particularly his critique of the vision-language-action (VLA) framework driving many of today’s embodied robots. Wang didn’t mince words: he called the VLA architecture “relatively dumb.” The main problem, in his view, is data, or the lack thereof. VLA models require vast, high-quality datasets to function effectively in the real world. While the scarcity of such data is widely acknowledged, many companies have pursued brute-force methods: gathering real-world robot data, generating simulation data, or building specialized data collection infrastructure. Wang believes that emphasis is misplaced. “People are payi
Describing VLA models as “relatively dumb,” he outlined an alternative approach to embodied intelligence. When most in the industry think of Unitree Robotics, they see a company focused on building robot hardware. But at the World Robot Conference (WRC) , founder Wang Xingxing offered a different narrative. During his keynote at WRC, Wang dedicated a large portion of his talk to large models, algorithms, and data. It was a shift that didn’t go unnoticed. His comments sparked debate, particularly his critique of the vision-language-action (VLA) framework driving many of today’s embodied robots.
saved by
related reading
- Robbyant - Exploring the Frontiers of Embodied Intelligence | 蚂蚁灵波科技 - 探索具身智能上限,打造物理世界的 AGI 平台technology.robbyant.com
- Sporks of AGIsergeylevine.substack.com
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- All Roads Lead to Robotics | Eric Jangevjang.com
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- Many Small Steps for Robots, One Giant Leap for Mankindnotboring.co
- Android Dreamsandroid-dreams.ai
- how we accidentally solved robotics by watching 1 million hours of YouTube – atharva's blogksagar.bearblog.dev
- The Embodied Internetben.bolte.cc
- Fully autonomous robots are much closer than you think – Sergey Levinedwarkesh.com
- How Claude Performs on Robotics Tasks \ Anthropicanthropic.com