A VLA with Open-World Generalization
Robots have come a long way over the past few years—they can perform impressive acrobatic feats, dance on stage, follow language commands and, in some of our own results, perform complex tasks like folding laundry or cleaning off a table. But the biggest challenge in robotics is not in performing feats of agility or dexterity, but generalization: the ability to figure out how to correctly perform even a simple task in a new setting or with new objects. Imagine a robot that needs to clean your home: every home is different, with different objects in different places. Generalization must occur at many levels. At the low level, the robot must understand how to pick up a spoon (by the handle) or plate (by the edge), even if it has not seen these specific spoons or plates before, and even if they are placed in a pile of dirty dishes. At a higher level, the robot must understand the semantics of each task—where to put clothes and shoes (ideally in the laundry hamper or closet, not on the bed
A VLA with Open-World Generalization π 0.5 : a VLA with Open-World Generalization Published April 22, 2025 Email research@physicalintelligence.company Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Manuel Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Allen Ren, Lucy Xiaoyang Shi, Laura Smith, Jost Tobias Springenberg, Kyle Stachowicz, James Tanner, Quan Vuong, Home
Explore this link on the map →saved by
related reading
- A VLA with Open-World Generalizationpi.website
- A Steerable Model with Emergent Capabilitiespi.website
- how we accidentally solved robotics by watching 1 million hours of YouTube – atharva's blogksagar.bearblog.dev
- State of Robot Learning, December 2025vedder.io
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- A VLA that Learns from Experiencepi.website
- Precise Manipulation with Efficient Online RLpi.website
- To Understand Language is to Understand Generalization | Eric Jangevjang.com
- RT-2: Vision-Language-Action Modelsrobotics-transformer2.github.io
- 𝜋₀: A Vision-Language-Action Flow Model for General Robot Controlarxiv.org