Video Generation Models Explosion 2024 - Yen-Chen Lin
yenchenlin.me · 4,798 words · saved by 2 readers
an example blog set up following the tutorial
Video generation models exploded onto the scene in 2024, sparked by the release of Sora from OpenAI. This blog post is my way of keeping track of the progress of this fascinating field. I will review all the key techniques that are used in building state-of-the-art video generation models (1). A comprehensive review of all text-to-image/text-to-video models is beyond the scope of this blog post. I will focus on research that has been published, productionized, or open-sourced. All of the videos and images are reproduced from the cited projects and papers, and the copyright belongs to the…
saved by
related reading
- Video models are zero-shot learners and reasonersarxiv.org
- [2404.02905] Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Predictionarxiv.org
- Can Video World Models Track Unobserved World States?joonghyuk.com
- What are Diffusion Models?lilianweng.github.io
- A Dive into Text-to-Video Modelshuggingface.co
- How do AI models generate videos? | MIT Technology Reviewtechnologyreview.com
- Seoul World Model: Grounding World Simulation Models in a Real-World Metropolisseoul-world-model.github.io
- What are Diffusion Models? | Lil'Loglilianweng.github.io
- Sand.aisand.ai
- Are Video Generation Models World Simulators? · Artificial Cognitionartificialcognition.net
- [2412.03603] HunyuanVideo: A Systematic Framework For Large Video Generative Modelsarxiv.org
- The First Fully General Computer Action Model | blogsi.inc