flâneur

RoboTTT: Context Scaling for Robot Policies

research.nvidia.com · 1,849 words · saved by 1 readers

RoboTTT integrates Test-Time Training into robot foundation models, scaling visuomotor context to 8K timesteps — three orders of magnitude beyond state-of-the-art policies — without growing inference latency.

Yunfan Jiang1,2, Yevgen Chebotar1, Ruijie Zheng1, Fengyuan Hu1, Yunhao Ge1, Jimmy Wu1, Tianyuan Dai1,3, Scott Reed1, Li Fei-Fei2,†, Yuke Zhu1,3,†, Linxi “Jim” Fan1,† 1NVIDIA2Stanford University3The University of Texas at Austin†Equal advising arXivPaper Abstract Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At…

saved by

related reading