State-of-the-Art Image Generative Models
I have aggregated some of the SotA image generative models released recently, with short summaries, visualizations and comments. The overall development is summarized, and the future trends are spe…
I have aggregated some of the SotA image generative models released recently, with short summaries, visualizations and comments. The overall development is summarized, and the future trends are speculated. Many of the statements and the results here are easily applicable to other non-textual modalities, such as audio and video. Summary: The papers we featured in this post belong to either of the following paradigms of SotA image generative models: VAE: VDVAE and VQVAE variants offer SotA diversity (NLL or recall). Furthermore, the sampling speed of VAEs without discrete bottleneck (e.g. VDVAE)
related reading
- What are Diffusion Models?lilianweng.github.io
- What are Diffusion Models? | Lil'Loglilianweng.github.io
- ⭐️ Diffusion Modelsandrewkchan.dev
- An In-Depth Analysis of VAR (NIPS 2024 Best Paper): Why Next-scale Prediciton Beats Diffusion Models?zhouyifan.net
- Yang Songyang-song.net
- [2404.02905] Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Predictionarxiv.org
- Generative modelling in latent space – Sander Dielemansander.ai
- Video Generation Models Explosion 2024yenchenlin.me
- Elucidating the Design Space of Diffusion-Based Generative Models | PDFarxiv.org
- Replicate - Run AI with an APIreplicate.com
- https://arxiv.org/pdf/2006.11239arxiv.org
- [2006.11239] Denoising Diffusion Probabilistic Modelsarxiv.org