Robbyant
Technology-driven and application-oriented. We build foundational large models for embodied AI: spatial perception (LingBot-Depth), VLA (LingBot-VLA), world models (LingBot-World), video action (LingBot-VA). Jointly embrace the new era of embodied intelligence. 技术驱动、场景导向,自研具身智能基础大模型,共迎具身智能新时代,共创幸福生活新场景。
Method Streaming 3D reconstruction is fundamentally a question of memory — what to keep, and in what form. LingBot-Map answers this with Geometric Context Attention (GCA), a small but structured streaming state that is learned end‑to‑end. GCA maintains three complementary contexts: an anchor for coordinate and scale grounding, a local pose‑reference window for dense local geometry, and a trajectory memory that compresses the full history into compact per‑frame tokens — keeping memory and compute per frame nearly constant on sequences of 10,000+ frames at ~20 FPS. Pipeline of LingBot-Map. A…
saved by
related reading
- Robbyant - Exploring the Frontiers of Embodied Intelligence | 蚂蚁灵波科技 - 探索具身智能上限,打造物理世界的 AGI 平台technology.robbyant.com
- GEN-1.5: Embodied Foundation Models are One-Shot Learners - Generalist AIgeneralistai.com
- Explore | alphaXivalphaxiv.org
- How Claude Performs on Robotics Tasks \ Anthropicanthropic.com
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- Physical Intelligence (π)pi.website
- The flavor of the bitter lesson for computer vision - Vincent Sitzmannvincentsitzmann.com
- A Steerable Model with Emergent Capabilitiespi.website
- Introducing Robostral Navigatemistral.ai
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- Explore | alphaXivalphaxiv.org
- GitHub - robotics-survey/Awesome-Robotics-Foundation-Modelsgithub.com