1X World Model | From Video to Action: A New Way Robots Learn
1x.tech · 2,162 words · saved by 1 readers
Home robots need common sense behavior and a deep understanding of the physical world.
Home robots need common sense behavior and a deep understanding of the physical world. Many robot foundation models today are vision-language-action models (VLAs), which take a pretrained VLM and add an output head to predict robot actions (PI0.6, Helix, Groot N1.5). VLMs benefit from internet-scale knowledge, but are trained on objectives that emphasize visual and semantic understanding over prediction of physical dynamics. Tens of thousands of hours of costly robot data are needed to teach a model how to solve tasks considered simple for a human. Additionally, auxiliary objectives are…
saved by
related reading
- Anirudha Majumdar (@Majumdar_Ani) on Xx.com
- Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AIrhoda.ai
- World Models: Computing the Uncomputablenotboring.co
- State of Robot Learning, December 2025vedder.io
- Generalist - GEN-1: Scaling Embodied Foundation Models to Masterygeneralistai.com
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- World Models | Rohit Bandarurohitbandaru.github.io
- The Model That Dreams the Worldmoe-capital.com
- Explore | alphaXivalphaxiv.org
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- A VLA with Open-World Generalizationpi.website
- How Claude Performs on Robotics Tasks \ Anthropicanthropic.com