2408.11039
arxiv.org · 7,407 words · saved by 1 readers
N/A
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model Chunting Zhouµ∗ Lili Yuµ∗ Arun Babuδ † Kushal Tirumalaµ Michihiro Yasunagaµ Leonid Shamisµ Jacob Kahnµ Xuezhe Maσ Luke Zettlemoyerµ Omer Levy† arXiv:2408.11039v1 [cs.AI] 20 Aug 2024…
related reading
- What are Diffusion Models? | Lil'Loglilianweng.github.io
- What are Diffusion Models?lilianweng.github.io
- The Illustrated Stable Diffusion – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Large Language Diffusion Modelsarxiv.org
- Scalable Diffusion Models with Transformersopenaccess.thecvf.com
- ⭐️ Diffusion Modelsandrewkchan.dev
- Kuleshov Group | How to Build a Diffusion Language Modelkuleshov-group.github.io
- Scalable Diffusion Models with Transformersarxiv.org
- 2403.09611.pdfarxiv.org
- Unified Multimodal Models as Auto-Encodersarxiv.org
- 2406.07524arxiv.org
- Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning | Researchai.meta.com