flâneur — a map of the web's best reading

Emergence of Human to Robot Transfer in Vision-Language-Action Models

physicalintelligence.company · 1,376 words · saved by 1 readers

One of the most exciting (and perhaps controversial) phenomena in large language models is emergence. As models and datasets become bigger, some capabilities, such as in-context learning and effective chain-of-thought reasoning, begin to appear only above a particular scale. One of the things that can emerge at scale with LLMs is the ability to more effectively leverage data, both through compositionality and generalization, and by utilizing other data sources, such as synthetic data produced via RL. As we scale up foundation models, they become generalists that can soak up diverse data sources in ways that smaller models cannot. In this post, we’ll discuss some of our recent results showing that transfer from human videos to robotic tasks emerges in robotic foundation models as we scale up the amount of robot training data. Based on this finding, we developed a method for using ego-centric data from humans to improve our models, providing a roughly 2x improvement on tasks where robot

Emergence of Human to Robot Transfer in Vision-Language-Action Models Emergence of Human to Robot Transfer in VLAs Published December 16, 2025 Email research@physicalintelligence.company Simar Kareer, Karl Pertsch, James Darpinian, Judy Hoffman, Danfei Xu, Sergey Levine, Chelsea Finn, Suraj Nair Paper human_to_robot.pdf Loading… Loading… Loading… Loading… Loading… Loading… Loading… Loading… Loading… Loading… Loading… Loading… Loading… Loading… Loading… Loading… HUMAN DATA DIVERSE AND LARGE SCALE ROBOT DATA NEW ROBOT CAPABILITIES Loading… Loading… Loading… Loading… Loading… Loading… Loading… Lo

Explore this link on the map →

related reading