[2509.02722] Planning with Reasoning using Vision Language World Model
Abstract:Effective planning requires strong world models, but high-level world models that can understand and reason about actions with semantic and temporal abstraction remain largely underdeveloped. We introduce the Vision Language World Model (VLWM), a foundation model trained for language-based world modeling on natural videos. Given visual observations, the VLWM first infers the overall goal achievements then predicts a trajectory composed of interleaved actions and world state changes. Those targets are extracted by iterative LLM Self-Refine conditioned on compressed future observations represented by Tree of Captions. The VLWM learns both an action policy and a dynamics model, which respectively facilitates reactive system-1 plan decoding and reflective system-2 planning via cost minimization. The cost evaluates the semantic distance between the hypothetical future states given by VLWM roll-outs and the expected goal state, and is measured by a critic model that we trained in a self-supervised manner. The VLWM achieves state-of-the-art Visual Planning for Assistance (VPA) performance on both benchmark evaluations and our proposed PlannerArena human evaluations, where system-2 improves the Elo score by +27% upon system-1. The VLWM models also outperforms strong VLM baselines on RoboVQA and WorldPrediction benchmark.
[2509.02722] Planning with Reasoning using Vision Language World Model --> Computer Science > Artificial Intelligence arXiv:2509.02722 (cs) [Submitted on 2 Sep 2025 ( v1 ), last revised 6 Sep 2025 (this version, v2)] Title: Planning with Reasoning using Vision Language World Model Authors: Delong Chen , Theo Moutakanni , Willy Chung , Yejin Bang , Ziwei Ji , Allen Bolourchi , Pascale Fung View a PDF of the paper titled Planning with Reasoning using Vision Language World Model, by Delong Chen and 6 other authors View PDF HTML (experimental) Abstract: Effective planning requires strong world mod
Explore this link on the map →saved by
related reading
- Explore | alphaXivalphaxiv.org
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- Semantic World Modelsweirdlabuw.github.io
- Semantic World Modelsarxiv.org
- World Models | Rohit Bandarurohitbandaru.github.io
- World Models: Computing the Uncomputablenotboring.co
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- pdfopenreview.net
- Video Language Planningvideo-language-planning.github.io
- A Functional Taxonomy of World Models - Dr. Fei-Fei Lidrfeifei.substack.com
- Language Models, World Models, and Human Model-Buildinglingo.csail.mit.edu
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixelsle-wm.github.io