Humanoid Locomotion as Next Token Prediction
We cast real-world humanoid control as a next token prediction problem, akin to predicting the next word in language. Our model is a causal transformer trained via autoregressive prediction. To account for the multi-modal nature of the data, we perform prediction in a modality-aligned way, and for each input token predict the next token from the same modality. This general formulation enables us to leverage data with missing modalities, like video trajectories without actions. We train our model on a collection of simulated trajectories coming from prior neural network policies, model-based controllers, motion capture data, and YouTube videos of humans. We show that our model enables a full-sized humanoid to walk in San Francisco zero-shot. Our model can transfer to the real world even when trained on only 27 hours of walking data, and can generalize to commands not seen during training like walking backward. These findings suggest a promising path toward learning challenging real-worl
Humanoid Locomotion as Next Token Prediction Marina Green Alta Plaza Park Balmy Alley Bay Bridge Bay Trail City Hall Ferry Building Market Street Harry Bridges Plaza Painted Ladies Palace of Fine Arts Crissy Field Crissy Field Beach Washington St California St Embarcadero Bart Humanoid Locomotion as Next Token Prediction Ilija Radosavovic Bike Zhang Baifeng Shi Jathushan Rajasegaran Sarthak Kamat Trevor Darrell Koushil Sreenath Jitendra Malik Ilija Radosavovic Bike Zhang Baifeng Shi Jathushan Rajasegaran Sarthak Kamat Trevor Darrell Koushil Sreenath Jitendra Malik University of California, Ber
Explore this link on the map →related reading
- State of Robot Learning, December 2025vedder.io
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AIrhoda.ai
- Explore | alphaXivalphaxiv.org
- Humanoid Atlas | Humanoid Robot Supply Chain Map, OEM Database & Industry Analysishumanoids.fyi
- A Steerable Model with Emergent Capabilitiespi.website
- e5b5c402bb7bd5e60bede6961d6fe39e-Paper-Conference.pdfproceedings.iclr.cc
- BeyondMimic: From Motion Tracking to Versatile Humanoid Control via Guided Diffusionarxiv.org
- A VLA with Open-World Generalizationpi.website
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixelsle-wm.github.io
- LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestionpku-epic.github.io
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robotsarxiv.org