flâneur

A Dive into Text-to-Video Models

huggingface.co · 1,962 words · saved by 1 readers

We’re on a journey to advance and democratize artificial intelligence through open source and open science.

Video samples generated with ModelScope. Text-to-video is next in line in the long list of incredible advances in generative models. As self-descriptive as it is, text-to-video is a fairly new computer vision task that involves generating a sequence of images from text descriptions that are both temporally and spatially consistent. While this task might seem extremely similar to text-to-image, it is notoriously more difficult. How do these models work, how do they differ from text-to-image models, and what kind of performance can we expect from them? In this blog post, we will discuss the…

saved by

related reading