✳flâneur — a map of the web's best reading
How can we get enough data to train a robot GPT?
itcanthink.substack.com · 2,145 words · saved by 1 readers
A thought experiment on scaling robot data collection to 2 trillion tokens
How can we get enough data to train a robot GPT? A thought experiment on scaling robot data collection to 2 trillion tokens Chris Paxton Jun 10, 2025 59 8 8 Share It’s no secret that large language models are trained on massive amounts of data - many trillions of tokens. Even the largest robot datasets are quite far from this; in a year, Physical Intelligence collected about 10,000 hours worth of robot data to train their first foundation model, PI0. Professor Ken Goldberg of UC Berkeley gave a talk which Andra Keay writes about on substack : the huge question of the “robot data gap.” A quick
Explore this link on the map →related reading
- State of Robot Learning, December 2025vedder.io
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- Pantograph: Building a Preschool for Robotspantograph.com
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- Many Small Steps for Robots, One Giant Leap for Mankindnotboring.co
- Android Dreamsandroid-dreams.ai
- The Scaling Hypothesis · Gwern.netgwern.net
- how we accidentally solved robotics by watching 1 million hours of YouTube – atharva's blogksagar.bearblog.dev
- GENE-26.5: Advancing Robotic Manipulation to Human Levelgenesis.ai
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robotsarxiv.org
- Generalist - GEN-1: Scaling Embodied Foundation Models to Masterygeneralistai.com
- Open X-Embodiment: Robotic Learning Datasets and RT-X Modelsarxiv.org