✳flâneur — a map of the web's best reading
RT-2: Vision-Language-Action Models
robotics-transformer2.github.io · 1,592 words · saved by 2 readers
Project page for RT-2
RT-2: Vision-Language-Action Models --> --> Your browser does not support the video tag. RT2: Vision-Language-Action Models RT-2 model picking up object given the prompt "pick up the extinct animal." RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control Anthony Brohan Noah Brown Justice Carbajal Yevgen Chebotar Xi Chen Krzysztof Choromanski Tianli Ding Danny Driess Avinava Dubey Chelsea Finn Pete Florence Chuyuan Fu Montse Gonzalez Arenas Keerthana Gopalakrishnan Kehang Han Karol Hausman Alex Herzog Jasmine Hsu Brian Ichter Alex Irpan Nikhil Joshi Ryan Julian Dmitry Kal
Explore this link on the map →saved by
related reading
- A VLA with Open-World Generalizationpi.website
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- how we accidentally solved robotics by watching 1 million hours of YouTube – atharva's blogksagar.bearblog.dev
- Helix: A Vision-Language-Action Model for Generalist Humanoid Controlfigure.ai
- Explore | alphaXivalphaxiv.org
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- A Steerable Model with Emergent Capabilitiespi.website
- [2602.10556] LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transferarxiv.org
- MolmoAct Action Reasoning Models that can Reason in Spacearxiv.org
- 45d74e190008c7bff2845ffc8e3facd3-Paper-Conference.pdfproceedings.iclr.cc
- The flavor of the bitter lesson for computer vision - Vincent Sitzmannvincentsitzmann.com