Vision-Language-Action (VLA) Models: A Review of Recent Progress — Xiangyu Li
xxxxyu.github.io · 974 words · saved by 1 readers
Recent VLAs are moving from discrete to continuous control and from single-system to dual-system designs.
I am new to this field — feel free to discuss and bring up any questions! This post is adapted from my slides. Background and Concepts The Concept of Vision-Language-Action (VLA) Models In my understanding, Vision-Language-Action (VLA) models1 are multimodal foundation models for embodied AI. They take vision (e.g., observations in video streams) and language (e.g., user instructions) as inputs, and generate low-level robot actions (i.e., the control policy) as outputs. A VLA uses a vision-language model (VLM) for vision-and-language-conditioned action generation. VLA concepts and the…
saved by
related reading
- [2510.13626] LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Modelsarxiv.org
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- Helix: A Vision-Language-Action Model for Generalist Humanoid Controlfigure.ai
- SimpleVLA-RL: Scaling VLA Training via Reinforcement Learningalphaxiv.org
- [2602.10098] VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Modelarxiv.org
- RT-2: Vision-Language-Action Modelsrobotics-transformer2.github.io
- A VLA with Open-World Generalizationpi.website
- How Claude Performs on Robotics Tasks \ Anthropicanthropic.com
- Moritz Reuss — Robotics & VLA Researchmbreuss.github.io
- Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Controlsteerable-policies.github.io
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website