How can we get enough data to train a robot GPT?
itcanthink.substack.com · 2,145 words · saved by 1 readers
A thought experiment on scaling robot data collection to 2 trillion tokens
How can we get enough data to train a robot GPT? A thought experiment on scaling robot data collection to 2 trillion tokens Chris Paxton Jun 10, 2025 59 8 8 Share It’s no secret that large language models are trained on massive amounts of data - many trillions of tokens. Even the largest robot datasets are quite far from this; in a year, Physical Intelligence collected about 10,000 hours worth of robot data to train their first foundation model, PI0. Professor Ken Goldberg of UC Berkeley gave a talk which Andra Keay writes about on substack : the huge question of the “robot data gap.” A quick
related reading
- Many Small Steps for Robots, One Giant Leap for Mankindnotboring.co
- State of Robot Learning, December 2025vedder.io
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- Sporks of AGIsergeylevine.substack.com
- All Roads Lead to Robotics | Eric Jangevjang.com
- Frontier of General-Purpose Robotics (2025)adampatni.com
- Nikolaus West (@NikolausWest) on Xx.com
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- Pantograph: Building a Preschool for Robotspantograph.com
- How Can We Make Robotics More like Generative Modeling? | Eric Jangevjang.com
- ACT-1: A Robot Foundation Model Trained on Zero Robot Data | Sunday Robotics | The helpful robotics companysunday.ai
- GitHub - adam-maj/robotics: A deep dive on the history of robotics and the future of humanoidsgithub.com