flâneur

Robbyant

technology.robbyant.com · 134 words · saved by 1 readers

Technology-driven and application-oriented. We build foundational large models for embodied AI: spatial perception (LingBot-Depth), VLA (LingBot-VLA), world models (LingBot-World), video action (LingBot-VA). Jointly embrace the new era of embodied intelligence. 技术驱动、场景导向,自研具身智能基础大模型,共迎具身智能新时代,共创幸福生活新场景。

Method Streaming 3D reconstruction is fundamentally a question of memory — what to keep, and in what form. LingBot-Map answers this with Geometric Context Attention (GCA), a small but structured streaming state that is learned end‑to‑end. GCA maintains three complementary contexts: an anchor for coordinate and scale grounding, a local pose‑reference window for dense local geometry, and a trajectory memory that compresses the full history into compact per‑frame tokens — keeping memory and compute per frame nearly constant on sequences of 10,000+ frames at ~20 FPS. Pipeline of LingBot-Map. A…

saved by

related reading