OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generation models, limiting their potential in real applications. In this paper, we propose OmniHuman, a Diffusion Transformer-based framework that scales up data by mixing motion-related conditions into the training phase. To this end, we introduce two training principles for these mixed conditions, along with the corresponding model architecture and inference strategy. These designs enable OmniHuman to fully leverage data-driven motion generation, ultimately achieving highly realistic human video generation. More importantly, OmniHuman supports
OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models Gaojie Lin ∗ Jianwen Jiang ∗† Jiaqi Yang ∗ Zerong Zheng ∗ Chao Liang ByteDance https://omnihuman-lab.github.io/ Abstract End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generation models, limiting their potential in real applications. In this paper, we propose OmniHuman, a Diffusion Transformer-based framework that scales up data by mixing motion-r
Explore this link on the map →saved by
related reading
- The First Fully General Computer Action Model | blogsi.inc
- What are Diffusion Models? | Lil'Loglilianweng.github.io
- Crossing the uncanny valley of conversational voice | Sesamesesame.com
- Replicate - Run AI with an APIreplicate.com
- ⭐️ Diffusion Modelsandrewkchan.dev
- How do AI models generate videos? | MIT Technology Reviewtechnologyreview.com
- Animate Anyonehumanaigc.github.io
- Woosh: A Sound Effects Foundation Modelarxiv.org
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robotsarxiv.org
- Yang Songyang-song.net
- Feature-wise transformationsdistill.pub