flâneur — a map of the web's best reading

How can we get enough data to train a robot GPT?

itcanthink.substack.com · 2,145 words · saved by 1 readers

A thought experiment on scaling robot data collection to 2 trillion tokens

How can we get enough data to train a robot GPT? A thought experiment on scaling robot data collection to 2 trillion tokens Chris Paxton Jun 10, 2025 59 8 8 Share It’s no secret that large language models are trained on massive amounts of data - many trillions of tokens. Even the largest robot datasets are quite far from this; in a year, Physical Intelligence collected about 10,000 hours worth of robot data to train their first foundation model, PI0. Professor Ken Goldberg of UC Berkeley gave a talk which Andra Keay writes about on substack : the huge question of the “robot data gap.” A quick

Explore this link on the map →

related reading