Video models are zero-shot learners and reasoners
arxiv.org · 15,189 words · saved by 2 readers
N/A
# link_13vn6wz50ev.pdf ## Metadata - PDFFormatVersion=1.7 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Thaddäus Wiedemer; Yuxuan Li; Paul Vicol; Shixiang Shane Gu; Nick Matarese; Kevin Swersky; Been Kim; Priyank Jaini; Robert Geirhos - Creator=arXiv GenPDF (tex2pdf:) - Custom.DOI=https://doi.org/10.48550/arXiv.2509.20328 - Custom.License=http://arxiv.org/licenses/nonexclusive-distrib/1.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.28 (TeX Live 2025) kpathsea version 6.4.1 - Custom
saved by
related reading
- Video Generation Models Explosion 2024yenchenlin.me
- Can Video World Models Track Unobserved World States?joonghyuk.com
- World Action Model Atlasjoeclinton.me
- [2602.10098] VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Modelarxiv.org
- [2404.02905] Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Predictionarxiv.org
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- The flavor of the bitter lesson for computer vision - Vincent Sitzmannvincentsitzmann.com
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- The First Fully General Computer Action Model | blogsi.inc
- A Dive into Text-to-Video Modelshuggingface.co
- [2411.02385] How Far is Video Generation from World Model: A Physical Law Perspectivearxiv.org
- Veo 3.1deepmind.google