Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AI
At Rhoda AI, we are building towards generalist robotics. Our Direct Video-Action Model (DVA) reformulates robot policies as video generation, unlocking data-efficient task learning, scaling, long-context memory, and one-shot learning.
Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AI Causal Video Models Are Data-Efficient Robot Policy Learners March 2026 · Rhoda AI Research At Rhoda AI, we are building towards generalist robotics. Our Direct Video-Action Model (DVA) reformulates robot policies as video generation, unlocking data-efficient task learning, scaling, long-context memory, and one-shot learning. Contents The Challenge of Generalist Robotics For decades, we have excelled at creating specialized robots — machines that perform a single, repetitive task with superhuman speed and accuracy in contr
saved by
related reading
- 1X World Model | From Video to Action: A New Way Robots Learn1x.tech
- 45d74e190008c7bff2845ffc8e3facd3-Paper-Conference.pdfproceedings.iclr.cc
- e5b5c402bb7bd5e60bede6961d6fe39e-Paper-Conference.pdfproceedings.iclr.cc
- [2603.08546] Interactive World Simulator for Robot Policy Training and Evaluationarxiv.org
- Anirudha Majumdar (@Majumdar_Ani) on Xx.com
- [2509.22407] EMMA: Generalizing Real-World Robot Manipulation via Generative Visual Transferarxiv.org
- [2506.09985] V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planningarxiv.org
- The First Fully General Computer Action Model | blogsi.inc
- State of Robot Learning, December 2025vedder.io
- [2206.11795] Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videosarxiv.org
- how we accidentally solved robotics by watching 1 million hours of YouTube – atharva's blogksagar.bearblog.dev
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website