Video Language Planning
We are interested in enabling visual planning for complex long-horizon tasks in the space of generated videos and language, leveraging recent advances in large generative models pretrained on Internet-scale data. To this end, we present video language planning (VLP), an algorithm that consists of a tree search procedure, where we train (i) vision-language models to serve as both policies and value functions, and (ii) text-to-video models as dynamics models. VLP takes as input a long-horizon task instruction and current image observation, and outputs a long video plan that provides detailed multimodal (video and language) specifications that describe how to complete the final task. VLP scales with increasing computation budget where more computation time results in improved video plans, and is able to synthesize long-horizon video plans across different robotics domains – from multi-object rearrangement, to multi-camera bi-arm dexterous manipulation. Generated video plans can be transla
Video Language Planning Video Language Planning Yilun Du Sherry Yang Pete Florence Fei Xia Ayzaan Wahid Brian Ichter Pierre Sermanet Tianhe Yu Pieter Abbeel Joshua B. Tenenbaum Leslie Kaelbling Andy Zeng Jonathan Tompson Google DeepMind MIT UC Berkeley Paper Code Data Results Abstract We are interested in enabling visual planning for complex long-horizon tasks in the space of generated videos and language, leveraging recent advances in large generative models pretrained on Internet-scale data. To this end, we present video language planning (VLP), an algorithm that consists of a tree search pr
Explore this link on the map →saved by
related reading
- [2509.02722] Planning with Reasoning using Vision Language World Modelarxiv.org
- The First Fully General Computer Action Model | blogsi.inc
- Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AIrhoda.ai
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- how we accidentally solved robotics by watching 1 million hours of YouTube – atharva's blogksagar.bearblog.dev
- Video models are zero-shot learners and reasonersarxiv.org
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- [2506.09985] V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planningarxiv.org
- 45d74e190008c7bff2845ffc8e3facd3-Paper-Conference.pdfproceedings.iclr.cc
- LLM Visualizationbbycroft.net
- [2201.07207] Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agentsarxiv.org
- Explore | alphaXivalphaxiv.org