Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blog
Quick glossary for readers new to VLA/WAM terminology VLA Vision-Language-Action model: a robot policy that starts from a pretrained VLM backbone and adapts it to generate actions from visual…
Quick glossary for readers new to VLA/WAM terminology VLA Vision-Language-Action model: a robot policy that starts from a pretrained VLM backbone and adapts it to generate actions from visual observations and language instructions. Large-scale VLM pretraining is a core part of the recipe. See Pi-0 and GR00T N1 . WAM World-Action Model: a policy that starts from a pretrained world-model or video backbone and adapts it to represent or predict how the scene changes over time and emit corresponding actions. We use WAM as the term throughout this post. VLM Vision-Language Model: a model pretrained
saved by
related reading
- General Instinct | Any frontier model. Any edge device.general-instinct.com
- [2602.10098] VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Modelarxiv.org
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- [2603.16666] Fast-WAM: Do World Action Models Need Test-time Future Imagination?arxiv.org
- World Action Model Atlasjoeclinton.me
- [2410.11758] Latent Action Pretraining from Videosarxiv.org
- Moritz Reuss — Robotics & VLA Researchmbreuss.github.io
- Learning to Act without Actionsarxiv.org
- What Matters for Latent Actions in Robot Learningcarldegio.github.io
- World Models | Rohit Bandarurohitbandaru.github.io
- 45d74e190008c7bff2845ffc8e3facd3-Paper-Conference.pdfproceedings.iclr.cc
- Robbyant - Exploring the Frontiers of Embodied Intelligence | 蚂蚁灵波科技 - 探索具身智能上限,打造物理世界的 AGI 平台technology.robbyant.com