flâneur

LingBot-VLA 2.0 Foundation Model - Robbyant

technology.robbyant.com · 407 words · saved by 1 readers

Technology-driven and application-oriented. We build foundational large models for embodied AI: vision foundation model (LingBot-Vision), spatial perception (LingBot-Depth), VLA (LingBot-VLA), world models (LingBot-World), video action (LingBot-VA), and video foundation models (LingBot-Video). Jointly embrace the new era of embodied intelligence. 技术驱动、场景导向,自研具身智能基础大模型,共迎具身智能新时代,共创幸福生活新场景。

LingBot-VLA 2.0 From Foundation to Application: Improving VLA Models in Practice Scaling up Pre‑training Dataset LingBot‑VLA 2.0 scales pre‑training with a broader data mixture: 50,000 hours of real robotic data and 10,000 hours of embodiment‑free egocentric manipulation data. Diverse Robot Source ~50000 hours Diverse Ego Source ~10000 hours Data processing pipeline The raw pool is split into robotic and egocentric streams, then filtered stage by stage. Gray bands represent rejected data; colored blocks are retained and passed forward. Pool Robotic Data Egocentric Data Smooth…

saved by

related reading