Does Scaling Web-Video Pre-training Help Real Robots Do Real Work? | Rhoda AI
rhoda.ai · 6,991 words · saved by 1 readers
We investigate scaling model size and pre-training compute for video-based robot policies, and benchmark on a real industrial manipulation task.
Video pre-training has recently become a cornerstone of robot foundation models, and with it, a widely held belief: scaling web-video pre-training leads to better downstream robot task performance. This belief has not been rigorously tested. Few studies examine the effect of scaling model size and compute, instead often substituting a binary comparison against a model trained from scratch. Fewer still have stepped beyond proxy metrics and simple lab tasks to investigate scaling on complex, long-horizon tasks representative of real-world use cases. We put this belief to the test with Direct…
saved by
related reading
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AIrhoda.ai
- The First Fully General Computer Action Model | blogsi.inc
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- 45d74e190008c7bff2845ffc8e3facd3-Paper-Conference.pdfproceedings.iclr.cc
- [2206.11795] Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videosarxiv.org
- 1X World Model | From Video to Action: A New Way Robots Learn1x.tech
- how we accidentally solved robotics by watching 1 million hours of YouTube – atharva's blogksagar.bearblog.dev
- [2506.09985] V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planningarxiv.org
- e5b5c402bb7bd5e60bede6961d6fe39e-Paper-Conference.pdfproceedings.iclr.cc
- Scaling Video Pretraining with Imagination Modelsinductionlabs.com