flâneur

Sand.ai - Advance AI to benefit everyone

sand.ai · 1,616 words · saved by 1 readers

Video generation is entering a new scaling stage. Longer duration, higher resolution, richer motion, synchronized audio, and stronger controllability all require more model capacity. For video diffusion, one of the core scaling pressures comes from sequence length. Compared with text, a video sample contains many more tokens: spatial patches across multiple frames, often combined with audio, text, reference images, or other conditioning signals. Under dense scaling, increasing model size means every token must pass through the full larger model. As video duration, resolution, and frame rate grow, this makes it increasingly difficult to scale model size aggressively in both the training side and inference side. Training stability is another constraint. Video diffusion models see different feature distributions across denoising timesteps; unified video/audio/text modeling introduces heterogeneous token types; and visual tokens are often highly correlated across space and time. As models

Video generation is entering a new scaling stage. Longer duration, higher resolution, richer motion, synchronized audio, and stronger controllability all require more model capacity. For video diffusion, one of the core scaling pressures comes from sequence length. Compared with text, a video sample contains many more tokens: spatial patches across multiple frames, often combined with audio, text, reference images, or other conditioning signals. Under dense scaling, increasing model size means every token must pass through the full larger model. As video duration, resolution, and frame rate…

saved by

related reading