Flexible Diffusion Modeling of Long Videos
We present a framework for video modeling based on denoising diffusion probabilistic models that produces long-duration video completions in a variety of realistic environments. We introduce a generative model that can at test-time sample any arbitrary subset of video frames conditioned on any other subset and present an architecture adapted for this purpose. Doing so allows us to efficiently compare and optimize a variety of schedules for the order in which frames in a long video are sampled and use selective sparse and long-range conditioning on previously sampled frames. We demonstrate improved video modeling over prior work on a number of datasets and sample temporally coherent videos over 25 minutes in length. We additionally release a new video modeling dataset and semantically meaningful metrics based on videos generated in the CARLA self-driving car simulator.
Flexible Diffusion Modeling of Long Videos William Harvey, Saeid Naderiparizi, Vaden Masrani, Christian Weilbach, Frank Wood∗ Department of Computer Science University of British Columbia Vancouver, Canada {wsgh,saeidnp,vadmas,weilbach,fwood}@cs.ubc.ca arXiv:2205.11495v3…
related reading
- Video Generation Models Explosion 2024yenchenlin.me
- The First Fully General Computer Action Model | blogsi.inc
- What are Diffusion Models? | Lil'Loglilianweng.github.io
- Sand.aisand.ai
- What are Diffusion Models?lilianweng.github.io
- A Dive into Text-to-Video Modelshuggingface.co
- How do AI models generate videos? | MIT Technology Reviewtechnologyreview.com
- ⭐️ Diffusion Modelsandrewkchan.dev
- [2006.11239] Denoising Diffusion Probabilistic Modelsarxiv.org
- Diffusion models from scratchchenyang.co
- MirageLSD: The First Live-Stream Diffusion AI Video Model | Decart AIabout.decart.ai
- 2409.02908arxiv.org