$π_0$: A Vision-Language-Action Flow Model for General Robot Control | alphaXiv
View recent discussion. Abstract: Robot learning holds tremendous promise to unlock the full potential of flexible, general, and dexterous robot systems, as well as to address some of the deepest questions in artificial intelligence. However, bringing robot learning to the level of generality required for effective real-world systems faces major obstacles in terms of data, generalization, and robustness. In this paper, we discuss how generalist robot policies (i.e., robot foundation models) can address these challenges, and how we can design effective generalist robot policies for complex and highly dexterous tasks. We propose a novel flow matching architecture built on top of a pre-trained vision-language model (VLM) to inherit Internet-scale semantic knowledge. We then discuss how this model can be trained on a large and diverse dataset from multiple dexterous robot platforms, including single-arm robots, dual-arm robots, and mobile manipulators. We evaluate our model in terms of its ability to perform tasks in zero shot after pre-training, follow language instructions from people and from a high-level VLM policy, and its ability to acquire new skills via fine-tuning. Our results cover a wide variety of tasks, such as laundry folding, table cleaning, and assembling boxes.
Submitted 08 Jan 2026 Abstract Robot learning holds tremendous promise to unlock the full potential of flexible, general, and dexterous robot systems, as well as to address some of the deepest questions in artificial intelligence. However, bringing robot learning to the level of generality required for effective real-world systems faces major obstacles in terms of data, generalization, and robustness. In this paper, we discuss how generalist robot policies (i.e., robot foundation models) can address these challenges, and how we can design effective generalist robot policies for complex and…
saved by
related reading
- SimpleVLA-RL: Scaling VLA Training via Reinforcement Learningalphaxiv.org
- Beyond Human Demonstrations: Diffusion-Based Reinforcement Learning to Generate Data for VLA Trainingalphaxiv.org
- 𝜋₀: A Vision-Language-Action Flow Model for General Robot Controlarxiv.org
- A Steerable Model with Emergent Capabilitiespi.website
- A VLA with Open-World Generalizationpi.website
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- RT-2: Vision-Language-Action Modelsrobotics-transformer2.github.io
- How Claude Performs on Robotics Tasks \ Anthropicanthropic.com
- Helix: A Vision-Language-Action Model for Generalist Humanoid Controlfigure.ai
- Physical Intelligence (π)pi.website
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AIrhoda.ai