LingBot-VLA 2.0 Foundation Model - Robbyant
Technology-driven and application-oriented. We build foundational large models for embodied AI: vision foundation model (LingBot-Vision), spatial perception (LingBot-Depth), VLA (LingBot-VLA), world models (LingBot-World), video action (LingBot-VA), and video foundation models (LingBot-Video). Jointly embrace the new era of embodied intelligence. 技术驱动、场景导向,自研具身智能基础大模型,共迎具身智能新时代,共创幸福生活新场景。
LingBot-VLA 2.0 From Foundation to Application: Improving VLA Models in Practice Scaling up Pre‑training Dataset LingBot‑VLA 2.0 scales pre‑training with a broader data mixture: 50,000 hours of real robotic data and 10,000 hours of embodiment‑free egocentric manipulation data. Diverse Robot Source ~50000 hours Diverse Ego Source ~10000 hours Data processing pipeline The raw pool is split into robotic and egocentric streams, then filtered stage by stage. Gray bands represent rejected data; colored blocks are retained and passed forward. Pool Robotic Data Egocentric Data Smooth…
saved by
related reading
- LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestionpku-epic.github.io
- [2602.12215] LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestionarxiv.org
- [2410.11758] Latent Action Pretraining from Videosarxiv.org
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- Robbyant - Exploring the Frontiers of Embodied Intelligence | 蚂蚁灵波科技 - 探索具身智能上限,打造物理世界的 AGI 平台technology.robbyant.com
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- Physical Intelligence (π)pi.website
- How Claude Performs on Robotics Tasks \ Anthropicanthropic.com
- Sporks of AGIsergeylevine.substack.com
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- A VLA with Open-World Generalizationpi.website