Understanding VQ-VAE (DALL-E Explained Pt. 1) - ML@B Blog
VQ-VAE is a powerful technique for learning discrete representations of complex data types like images, video, or audio. This technique has played a key role in recent state of the art works like OpenAI's DALL-E and Jukebox models.
ML@B Blog Machine Learning at Berkeley is a student organization at UC Berkeley By Machine Learning · Over 4,000 subscribers By subscribing, you agree Substack's Terms of Use, and acknowledge its Information Collection Notice and Privacy Policy. ML@B Blog Home Archive About DiffusionGemma Explained by Timothy Gao Aug 10 • Machine Learning at Berkeley Benchmarking Biology’s AI Agent: ML@B's Collaboration with LatchBio By: Lucas Gu, Daniel Grant, Rishi Athavale, Maggie Dong Apr 15 • Machine Learning at Berkeley The Prose of Proteins - A Lesson in Taste and Vision through the…
saved by
related reading
- What are Diffusion Models?lilianweng.github.io
- kevin frans blogkvfrans.com
- Diffusion Models as a kind of VAE | Angus Turnerangusturner.github.io
- Video Generation Models Explosion 2024yenchenlin.me
- GenAI Handbookgenai-handbook.github.io
- Trending Papers - Hugging Facepaperswithcode.com
- Understanding VQ-VAE (DALL-E Explained Pt. 1)mlberkeley.substack.com
- David Duvenaudcs.toronto.edu
- Generative modelling in latent space – Sander Dielemansander.ai
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- Explore | alphaXivalphaxiv.org
- The 2025 AI Engineering Reading List - Latent.Spacelatent.space