State-of-the-Art Image Generative Models
I have aggregated some of the SotA image generative models released recently, with short summaries, visualizations and comments. The overall development is summarized, and the future trends are spe…
I have aggregated some of the SotA image generative models released recently, with short summaries, visualizations and comments. The overall development is summarized, and the future trends are speculated. Many of the statements and the results here are easily applicable to other non-textual modalities, such as audio and video. Summary: The papers we featured in this post belong to either of the following paradigms of SotA image generative models: VAE: VDVAE and VQVAE variants offer SotA diversity (NLL or recall). Furthermore, the sampling speed of VAEs without discrete bottleneck (e.g. VDVAE)
Explore this link on the map →related reading
- What are Diffusion Models? | Lil'Loglilianweng.github.io
- ⭐️ Diffusion Modelsandrewkchan.dev
- Yang Songyang-song.net
- Generative modelling in latent space – Sander Dielemansander.ai
- Replicate - Run AI with an APIreplicate.com
- Diffusion Models in AI – Everything You Need to Know – Unite.AIunite.ai
- Bare-bones Diffusion Modelsmadebyoll.in
- [2605.05331] ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parametersarxiv.org
- [2105.05233] Diffusion Models Beat GANs on Image Synthesisarxiv.org
- The Illustrated Stable Diffusion – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Diffusion models from scratchchenyang.co
- [2006.11239] Denoising Diffusion Probabilistic Modelsarxiv.org