Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AI
At Rhoda AI, we are building towards generalist robotics. Our Direct Video-Action Model (DVA) reformulates robot policies as video generation, unlocking data-efficient task learning, scaling, long-context memory, and one-shot learning.
Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AI Causal Video Models Are Data-Efficient Robot Policy Learners March 2026 · Rhoda AI Research At Rhoda AI, we are building towards generalist robotics. Our Direct Video-Action Model (DVA) reformulates robot policies as video generation, unlocking data-efficient task learning, scaling, long-context memory, and one-shot learning. Contents The Challenge of Generalist Robotics For decades, we have excelled at creating specialized robots — machines that perform a single, repetitive task with superhuman speed and accuracy in contr
Explore this link on the map →saved by
related reading
- 45d74e190008c7bff2845ffc8e3facd3-Paper-Conference.pdfproceedings.iclr.cc
- e5b5c402bb7bd5e60bede6961d6fe39e-Paper-Conference.pdfproceedings.iclr.cc
- [2603.08546] Interactive World Simulator for Robot Policy Training and Evaluationarxiv.org
- [2509.22407] EMMA: Generalizing Real-World Robot Manipulation via Generative Visual Transferarxiv.org
- The First Fully General Computer Action Model | blogsi.inc
- State of Robot Learning, December 2025vedder.io
- how we accidentally solved robotics by watching 1 million hours of YouTube – atharva's blogksagar.bearblog.dev
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- [2602.10556] LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transferarxiv.org
- Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Controlsteerable-policies.github.io
- Generalist - GEN-1: Scaling Embodied Foundation Models to Masterygeneralistai.com
- [2506.09985] V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planningarxiv.org