Emergence of Human to Robot Transfer in Vision-Language-Action Models
One of the most exciting (and perhaps controversial) phenomena in large language models is emergence. As models and datasets become bigger, some capabilities, such as in-context learning and effective chain-of-thought reasoning, begin to appear only above a particular scale. One of the things that can emerge at scale with LLMs is the ability to more effectively leverage data, both through compositionality and generalization, and by utilizing other data sources, such as synthetic data produced via RL. As we scale up foundation models, they become generalists that can soak up diverse data sources in ways that smaller models cannot. In this post, we’ll discuss some of our recent results showing that transfer from human videos to robotic tasks emerges in robotic foundation models as we scale up the amount of robot training data. Based on this finding, we developed a method for using ego-centric data from humans to improve our models, providing a roughly 2x improvement on tasks where robot
Emergence of Human to Robot Transfer in Vision-Language-Action Models Emergence of Human to Robot Transfer in VLAs Published December 16, 2025 Email research@physicalintelligence.company Simar Kareer, Karl Pertsch, James Darpinian, Judy Hoffman, Danfei Xu, Sergey Levine, Chelsea Finn, Suraj Nair Paper human_to_robot.pdf Loading… Loading… Loading… Loading… Loading… Loading… Loading… Loading… Loading… Loading… Loading… Loading… Loading… Loading… Loading… Loading… HUMAN DATA DIVERSE AND LARGE SCALE ROBOT DATA NEW ROBOT CAPABILITIES Loading… Loading… Loading… Loading… Loading… Loading… Loading… Lo
Explore this link on the map →related reading
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- how we accidentally solved robotics by watching 1 million hours of YouTube – atharva's blogksagar.bearblog.dev
- [2602.10556] LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transferarxiv.org
- EgoScaleresearch.nvidia.com
- A Steerable Model with Emergent Capabilitiespi.website
- RT-2: Vision-Language-Action Modelsrobotics-transformer2.github.io
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- 45d74e190008c7bff2845ffc8e3facd3-Paper-Conference.pdfproceedings.iclr.cc
- [2509.22407] EMMA: Generalizing Real-World Robot Manipulation via Generative Visual Transferarxiv.org
- A VLA with Open-World Generalizationpi.website
- Fully autonomous robots are much closer than you think – Sergey Levinedwarkesh.com