Sand.ai - Advance AI to benefit everyone
Video generation is entering a new scaling stage. Longer duration, higher resolution, richer motion, synchronized audio, and stronger controllability all require more model capacity. For video diffusion, one of the core scaling pressures comes from sequence length. Compared with text, a video sample contains many more tokens: spatial patches across multiple frames, often combined with audio, text, reference images, or other conditioning signals. Under dense scaling, increasing model size means every token must pass through the full larger model. As video duration, resolution, and frame rate grow, this makes it increasingly difficult to scale model size aggressively in both the training side and inference side. Training stability is another constraint. Video diffusion models see different feature distributions across denoising timesteps; unified video/audio/text modeling introduces heterogeneous token types; and visual tokens are often highly correlated across space and time. As models
Video generation is entering a new scaling stage. Longer duration, higher resolution, richer motion, synchronized audio, and stronger controllability all require more model capacity. For video diffusion, one of the core scaling pressures comes from sequence length. Compared with text, a video sample contains many more tokens: spatial patches across multiple frames, often combined with audio, text, reference images, or other conditioning signals. Under dense scaling, increasing model size means every token must pass through the full larger model. As video duration, resolution, and frame rate…
saved by
related reading
- Home | Genmogenmo.ai
- Video Generation Models Explosion 2024yenchenlin.me
- The First Fully General Computer Action Model | blogsi.inc
- A Dive into Text-to-Video Modelshuggingface.co
- Flexible Diffusion Modeling of Long Videosarxiv.org
- How do AI models generate videos? | MIT Technology Reviewtechnologyreview.com
- What are Diffusion Models? | Lil'Loglilianweng.github.io
- What are Diffusion Models?lilianweng.github.io
- francesco215.github.io/autoregressive_diffusion/francesco215.github.io
- ⭐️ Diffusion Modelsandrewkchan.dev
- MirageLSD: The First Live-Stream Diffusion AI Video Model | Decart AIabout.decart.ai
- OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Modelsarxiv.org