[2607.15275] RoboTTT: Context Scaling for Robot Policies
Abstract:Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturbations, and stronger performance on multi-stage, long-horizon tasks. We also observe, for the first time, steady gains in closed-loop performance as pretraining context length scales. At its core, RoboTTT integrates Test-Time Training into robot foundation models such as Vision-Language-Action policies, yielding a sequence model whose recurrent state consists of fast weights, parameters updated by gradient descent during both training and inference, compressing histories into weight space and retrieving contextual information for long-context conditioning. To scale training context length, the recipe combines sequence action forcing with truncated backpropagation through time. On challenging real-robot manipulation tasks, RoboTTT improves overall performance by 87% over the single-step context baseline and fully completes a five-minute, ten-stage assembly task, which no baseline ever does. RoboTTT trained with 8K-timestep context outperforms the same model pretrained with 1K timesteps by 62%, suggesting context length as a new scaling axis for robot foundation models. Videos are available at this https URL
2026-7-17 RoboTTT: Context Scaling for Robot Policies Yunfan Jiang1,2 , Yevgen Chebotar1 , Ruijie Zheng1 , Fengyuan Hu1 , Yunhao Ge1 Jimmy Wu1 , Tianyuan Dai1,3 , Scott Reed1 , Li Fei-Fei2,† , Yuke Zhu1,3,† , Linxi “Jim” Fan1,† 1 NVIDIA 2 Stanford University 3 The University of Texas at Austin † Equal advising research.nvidia.com/labs/gear/robottt…
saved by
related reading
- Introducing S1: In-Context Learning for Roboticsskild.ai
- RoboTTT: Context Scaling for Robot Policiesresearch.nvidia.com
- Moritz Reuss — Robotics & VLA Researchmbreuss.github.io
- State of Robot Learning, December 2025vedder.io
- ACT-1: A Robot Foundation Model Trained on Zero Robot Data | Sunday Robotics | The helpful robotics companysunday.ai
- GEN-1.5: Embodied Foundation Models are One-Shot Learners - Generalist AIgeneralistai.com
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AIrhoda.ai
- How Claude Performs on Robotics Tasks \ Anthropicanthropic.com
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- A Steerable Model with Emergent Capabilitiespi.website
- Explore | alphaXivalphaxiv.org