[2602.10556] LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer
Abstract:A long-standing goal in robotics is a generalist policy that can be deployed zero-shot on new robot embodiments without per-embodiment adaptation. Despite large-scale multi-embodiment pre-training, existing Vision-Language-Action models (VLAs) remain tightly coupled to their training embodiments and typically require costly fine-tuning. We introduce Language-Action Pre-training (LAP), a simple recipe that represents low-level robot actions directly in natural language, aligning action supervision with the pre-trained vision-language model's input-output distribution. LAP requires no learned tokenizer, no costly annotation, and no embodiment-specific architectural design. Based on LAP, we present LAP-3B, which to the best of our knowledge is the first VLA to achieve substantial zero-shot transfer to previously unseen robot embodiments without any embodiment-specific fine-tuning. Across multiple novel robots and manipulation tasks, LAP-3B attains over 50% average zero-shot success, delivering roughly a 2x improvement over the strongest prior VLAs. We further show that LAP enables efficient adaptation and favorable scaling, while unifying action prediction and VQA in a shared language-action format that yields additional gains through co-training.
# link_343osg61ay.pdf ## Metadata - PDFFormatVersion=1.7 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Lihan Zha; Asher J. Hancock; Mingtong Zhang; Tenny Yin; Yixuan Huang; Dhruv Shah; Allen Z. Ren; Anirudha Majumdar - Creator=arXiv GenPDF (tex2pdf:57610bf) - Custom.DOI=https://doi.org/10.48550/arXiv.2602.10556 - Custom.License=http://creativecommons.org/licenses/by/4.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.28 (TeX Live 2025) kpathsea version 6.4.1 - Custom.arXivID=https://ar
Explore this link on the map →saved by
related reading
- 45d74e190008c7bff2845ffc8e3facd3-Paper-Conference.pdfproceedings.iclr.cc
- [2509.22407] EMMA: Generalizing Real-World Robot Manipulation via Generative Visual Transferarxiv.org
- e5b5c402bb7bd5e60bede6961d6fe39e-Paper-Conference.pdfproceedings.iclr.cc
- Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Modelsarxiv.org
- [2603.08546] Interactive World Simulator for Robot Policy Training and Evaluationarxiv.org
- Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AIrhoda.ai
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- Precise Manipulation with Efficient Online RLpi.website
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- [2410.11758] Latent Action Pretraining from Videosarxiv.org
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- A VLA with Open-World Generalizationpi.website