[2602.10098] VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
Abstract:Pretraining Vision-Language-Action (VLA) policies on internet-scale video is appealing, yet current latent-action objectives often learn the wrong thing: they remain anchored to pixel variation rather than action-relevant state transitions, making them vulnerable to appearance bias, nuisance motion, and information leakage. We introduce VLA-JEPA, a JEPA-style pretraining framework that sidesteps these pitfalls by design. The key idea is leakage-free state prediction: a target encoder produces latent representations from future frames, while the student pathway sees only the current observation -- future information is used solely as supervision targets, never as input. By predicting in latent space rather than pixel space, VLA-JEPA learns dynamics abstractions that are robust to camera motion and irrelevant background changes. This yields a simple two-stage recipe -- JEPA pretraining followed by action-head fine-tuning -- without the multi-stage complexity of prior latent-action pipelines. Experiments on LIBERO, LIBERO-Plus, SimplerEnv and real-world manipulation tasks show that VLA-JEPA achieves consistent gains in generalization and robustness over existing methods.
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model Jingwen Sun1,2,∗ , Wenyao Zhang3,5∗ , Zekun Qi4 , Shaojie Ren2,6 , Zezhi Liu2,7 , Hanxin Zhu1 , Guangzhong Sun1 , Xin Jin2,5,† , Zhibo Chen1,2,† 1 University of Science and Technology of China 2 Zhongguancun Academy, Beijing, China 3 Shanghai Jiao Tong University 4 Tsinghua…
saved by
related reading
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- World Action Model Atlasjoeclinton.me
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- What Matters for Latent Actions in Robot Learningcarldegio.github.io
- Moritz Reuss — Robotics & VLA Researchmbreuss.github.io
- Can Video World Models Track Unobserved World States?joonghyuk.com
- [2410.11758] Latent Action Pretraining from Videosarxiv.org
- The flavor of the bitter lesson for computer vision - Vincent Sitzmannvincentsitzmann.com
- 45d74e190008c7bff2845ffc8e3facd3-Paper-Conference.pdfproceedings.iclr.cc
- [2506.09985] V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planningarxiv.org
- how we accidentally solved robotics by watching 1 million hours of YouTube – atharva's blogksagar.bearblog.dev
- Vision-Language-Action (VLA) Models: A Review of Recent Progressxxxxyu.github.io