✳flâneur — a map of the web's best reading
Video models are zero-shot learners and reasoners
arxiv.org · 15,189 words · saved by 2 readers
N/A
# link_13vn6wz50ev.pdf ## Metadata - PDFFormatVersion=1.7 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Thaddäus Wiedemer; Yuxuan Li; Paul Vicol; Shixiang Shane Gu; Nick Matarese; Kevin Swersky; Been Kim; Priyank Jaini; Robert Geirhos - Creator=arXiv GenPDF (tex2pdf:) - Custom.DOI=https://doi.org/10.48550/arXiv.2509.20328 - Custom.License=http://arxiv.org/licenses/nonexclusive-distrib/1.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.28 (TeX Live 2025) kpathsea version 6.4.1 - Custom
Explore this link on the map →saved by
related reading
- Seoul World Model: Grounding World Simulation Models in a Real-World Metropolisseoul-world-model.github.io
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- The flavor of the bitter lesson for computer vision - Vincent Sitzmannvincentsitzmann.com
- The First Fully General Computer Action Model | blogsi.inc
- Explore | alphaXivalphaxiv.org
- Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AIrhoda.ai
- The Model That Dreams the Worldmoe-capital.com
- DeepSeek-R1arxiv.org
- [2506.09985] V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planningarxiv.org
- [2411.02385] How Far is Video Generation from World Model: A Physical Law Perspectivearxiv.org
- Are Video Generation Models World Simulators? · Artificial Cognitionartificialcognition.net