[2602.12215] LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion
Abstract:Recent robot foundation models largely rely on large-scale behavior cloning, which imitates expert actions but discards transferable dynamics knowledge embedded in heterogeneous embodied data. While the Unified World Model (UWM) formulation has the potential to leverage such diverse data, existing instantiations struggle to scale to foundation-level due to coarse data usage and fragmented datasets. We introduce LDA-1B, a robot foundation model that scales through universal embodied data ingestion by jointly learning dynamics, policy, and visual forecasting, assigning distinct roles to data of varying quality. To support this regime at scale, we assemble and standardize EI-30k, an embodied interaction dataset comprising over 30k hours of human and robot trajectories in a unified format. Scalable dynamics learning over such heterogeneous data is enabled by prediction in a structured DINO latent space, which avoids redundant pixel-space appearance modeling. Complementing this representation, LDA-1B employs a multi-modal diffusion transformer to handle asynchronous vision and action streams, enabling stable training at the 1B-parameter scale. Experiments in simulation and the real world show LDA-1B outperforms prior methods (e.g., $\pi_{0.5}$) by up to 21\%, 48\%, and 23\% on contact-rich, dexterous, and long-horizon tasks, respectively. Notably, LDA-1B enables data-efficient fine-tuning, gaining 10\% by leveraging 30\% low-quality trajectories typically harmful and discarded.
[2602.12215] LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion --> Computer Science > Robotics arXiv:2602.12215 (cs) [Submitted on 12 Feb 2026 ( v1 ), last revised 3 Jun 2026 (this version, v2)] Title: LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion Authors: Jiangran Lyu , Kai Liu , Xuheng Zhang , Haoran Liao , Yusen Feng , Wenxuan Zhu , Tingrui Shen , Jiayi Chen , Jiazhao Zhang , Yifei Dong , Wenbo Cui , Senmao Qi , Shuo Wang , Yixin Zheng , Mi Yan , Xuesong Shi , Haoran Li , Dongbin Zhao , Ming-Yu Liu , Zhizheng Zhang , Li Y
saved by
related reading
- LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestionpku-epic.github.io
- Robbyant - Exploring the Frontiers of Embodied Intelligence | 蚂蚁灵波科技 - 探索具身智能上限,打造物理世界的 AGI 平台technology.robbyant.com
- Learning to Act without Actionsarxiv.org
- [2410.11758] Latent Action Pretraining from Videosarxiv.org
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- GEN-1.5: Embodied Foundation Models are One-Shot Learners - Generalist AIgeneralistai.com
- State of Robot Learning, December 2025vedder.io
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- Sporks of AGIsergeylevine.substack.com
- A Steerable Model with Emergent Capabilitiespi.website
- 45d74e190008c7bff2845ffc8e3facd3-Paper-Conference.pdfproceedings.iclr.cc