An In-Depth Analysis of VAR (NIPS 2024 Best Paper): Why Next-scale Prediciton Beats Diffusion Models? | Yifan Zhou's Blog
This past April, Peking University (PKU) and ByteDance published a paper on Arxiv titled Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction, introducing a brand-new pa
This past April, Peking University (PKU) and ByteDance published a paper on Arxiv titled Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction, introducing a brand-new paradigm for image generation called Visual Autoregressive Modeling (VAR). This autoregressive generation method represents high-definition images as multi-scale token images and replaces the previously popular next-token prediction with a next-scale prediction approach. On the ImageNet $256×256$ image generation task, VAR outperforms DiT. Our research group was quick to read the paper and found…
saved by
related reading
- [2404.02905] Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Predictionarxiv.org
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- What are Diffusion Models? | Lil'Loglilianweng.github.io
- Video Generation Models Explosion 2024yenchenlin.me
- What are Diffusion Models?lilianweng.github.io
- francesco215.github.io/autoregressive_diffusion/francesco215.github.io
- [2605.05331] ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parametersarxiv.org
- FlexTok: Resampling Images into 1D Token Sequences of Flexible Lengtharxiv.org
- State-of-the-Art Image Generative Models – Aran Komatsuzakiarankomatsuzaki.wordpress.com
- Generative modelling in latent space – Sander Dielemansander.ai
- Yang Songyang-song.net
- https://arxiv.org/pdf/2006.11239arxiv.org