Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blog
Quick glossary for readers new to VLA/WAM terminology VLA Vision-Language-Action model: a robot policy that starts from a pretrained VLM backbone and adapts it to generate actions from visual…
Quick glossary for readers new to VLA/WAM terminology VLA Vision-Language-Action model: a robot policy that starts from a pretrained VLM backbone and adapts it to generate actions from visual observations and language instructions. Large-scale VLM pretraining is a core part of the recipe. See Pi-0 and GR00T N1 . WAM World-Action Model: a policy that starts from a pretrained world-model or video backbone and adapts it to represent or predict how the scene changes over time and emit corresponding actions. We use WAM as the term throughout this post. VLM Vision-Language Model: a model pretrained
Explore this link on the map →saved by
related reading
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- 45d74e190008c7bff2845ffc8e3facd3-Paper-Conference.pdfproceedings.iclr.cc
- Video models are zero-shot learners and reasonersarxiv.org
- The flavor of the bitter lesson for computer vision - Vincent Sitzmannvincentsitzmann.com
- [2603.16666] Fast-WAM: Do World Action Models Need Test-time Future Imagination?arxiv.org
- World Models | Rohit Bandarurohitbandaru.github.io
- [2410.11758] Latent Action Pretraining from Videosarxiv.org
- [2509.02722] Planning with Reasoning using Vision Language World Modelarxiv.org
- [2602.10556] LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transferarxiv.org
- The Model That Dreams the Worldmoe-capital.com
- A VLA with Open-World Generalizationpi.website
- 𝜋₀: A Vision-Language-Action Flow Model for General Robot Controlarxiv.org