State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reuss
ICLR’s open submission policy gives a rare, real-time view of what the community is actually building. This post distills the state of Vision-Language-Action (VLA) Models research slice of ICLR 2026: what ‘counts’ as a VLA (and why that definition matters), what people are working on in the VLA field (discrete diffusion, embodied reasoning, new tokenizers), how to read benchmark results in VLA research, and the not-so-invisible frontier gap that sim leaderboards hide. Each autumn, the ICLR publicly releases all anonymous submissions a few weeks after the deadline, providing a unique real-time snapshot of ongoing research around the world without the typical six-month delay of other top ML conferences. Given my personal research interests, I wanted to analyze Vision-Language-Action (VLA) Models research and share insights from this year's submissions. In this blog post, I briefly explain what VLAs are, share my findings about current trends and challenges in VLA research, highlighting
State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reuss State of VLA Research at ICLR 2026 October 2025 • by Moritz Reuss ICLR’s open submission policy gives a rare, real-time view of what the community is actually building. This post distills the state of Vision-Language-Action (VLA) Models research slice of ICLR 2026: what ‘counts’ as a VLA (and why that definition matters), what people are working on in the VLA field (discrete diffusion, embodied reasoning, new tokenizers), how to read benchmark results in VLA research, and the not-so-invisible frontier gap that sim leade
Explore this link on the map →saved by
related reading
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- 45d74e190008c7bff2845ffc8e3facd3-Paper-Conference.pdfproceedings.iclr.cc
- A VLA with Open-World Generalizationpi.website
- The flavor of the bitter lesson for computer vision - Vincent Sitzmannvincentsitzmann.com
- Video models are zero-shot learners and reasonersarxiv.org
- RT-2: Vision-Language-Action Modelsrobotics-transformer2.github.io
- [2410.11758] Latent Action Pretraining from Videosarxiv.org
- Explore | alphaXivalphaxiv.org
- [2602.10556] LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transferarxiv.org
- pistar06.pdfpi.website
- Precise Manipulation with Efficient Online RLpi.website