Show-Harness
showlab.github.io · 1,741 words · saved by 1 readers
Show-Harness: Just a VLM Agent Can Play Robots
Just a VLM Agent Can Play Robots Yanzhe Chen*, Zechen Bai*, Zhijun Cao*, Wenzheng Zeng*, Kevin Qinghong Lin, Yiqi Lin, Guoqiang Liang, Kevin Yuchen Ma, Qiming Huang, Mike Zheng Shou† Show Lab, National University of Singapore *Equal contribution · †Corresponding author Paper Daily Paper GitHub Models Dataset One interface across scenes, tasks and embodiments. Frontier VLMs hold the top band at seconds per step, and the same interface fine-tuned onto open-source backbones holds it from 12 Hz up to 33 Hz — while π0.5 and GR00T sit a band below at 20 and 24 Hz. Abstract Foundation…
saved by
related reading
- Robot-use agentsweb.mit.edu
- How Claude Performs on Robotics Tasks \ Anthropicanthropic.com
- Helix: A Vision-Language-Action Model for Generalist Humanoid Controlfigure.ai
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Controlsteerable-policies.github.io
- Sporks of AGIsergeylevine.substack.com
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- Explore | alphaXivalphaxiv.org
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- Vision-Language-Action (VLA) Models: A Review of Recent Progressxxxxyu.github.io
- Robbyant - Exploring the Frontiers of Embodied Intelligence | 蚂蚁灵波科技 - 探索具身智能上限,打造物理世界的 AGI 平台technology.robbyant.com