Unified Multimodal Models as Auto-Encoders
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. Image-to-text (I2T) understanding and text-to-image (T2I) generation are two fundamental, important yet traditionally isolated multimodal tasks. Despite their intrinsic connection, existing approaches typically optimize them independently, missing the opportunity for mutual enhancement. In this paper, we argue that the both tasks can be connected under a shared Auto-Encoder perspective, where text serves as the intermediate latent representation bridging the two directions — encoding images into textual semantics (I2T) and decoding text back into images (T2I). Our key insight is that if the encoder truly “understands” the image, it should capture all essential structure, and if the decoder truly “understands” the text, it should recover that structure fa
Unified Multimodal Models as Auto-Encoders Zhiyuan Yan 1,2 ⋄,⋆ , Kaiqing Lin 1 ⋄ , Zongjian Li 1,3 ⋄ , Junyan Ye 4 ⋄ , Hui Han 1 , Haochen Wang 2,6 ⋆ , Zhendong Wang 5 , Bin Lin 1,3 , Hao Li 1 , Xinyan Xiao 2 , Jingdong Wang 2 , Haifeng Wang 2 , Li Yuan 1 † 1 Shenzhen Graduate School, Peking University 2 Baidu, 3 Rabbitpre AI 4 SYSU, 5 USTC, 6 CASIA zhiyuanyan@stu.pku.edu.cn Abstract † † ⋄ \diamond Equal Contribution, ⋆ \star Work done during an internship at Baidu Star Program, † Corresponding Author Image-to-text (I2T) understanding and text-to-image (T2I) generation are two fundamental, imp
Explore this link on the map →saved by
related reading
- [2301.13823] Grounding Language Models to Images for Multimodal Generationarxiv.org
- The Illustrated Stable Diffusion – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- What are Diffusion Models? | Lil'Loglilianweng.github.io
- 2403.09611.pdfarxiv.org
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversiontextual-inversion.github.io
- Introducing CM3leon, a more efficient, state-of-the-art generative model for text and imagesai.meta.com
- ImageBind: Holistic AI learning across six modalitiesai.facebook.com
- Generative modelling in latent space – Sander Dielemansander.ai
- ⭐️ Diffusion Modelsandrewkchan.dev
- Editing Text in Images with AI | Towards Data Sciencetowardsdatascience.com
- Yang Songyang-song.net
- cs.unc.edu/~mbansal/teaching/nlp-comp790-590-spring23.htmlcs.unc.edu