flâneur

Vision-Language-Action (VLA) Models: A Review of Recent Progress — Xiangyu Li

xxxxyu.github.io · 974 words · saved by 1 readers

Recent VLAs are moving from discrete to continuous control and from single-system to dual-system designs.

I am new to this field — feel free to discuss and bring up any questions! This post is adapted from my slides. Background and Concepts The Concept of Vision-Language-Action (VLA) Models In my understanding, Vision-Language-Action (VLA) models1 are multimodal foundation models for embodied AI. They take vision (e.g., observations in video streams) and language (e.g., user instructions) as inputs, and generate low-level robot actions (i.e., the control policy) as outputs. A VLA uses a vision-language model (VLM) for vision-and-language-conditioned action generation. VLA concepts and the…

saved by

related reading