flâneur — a map of the web's best reading

Unified Multimodal Models as Auto-Encoders

arxiv.org · 14,239 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. Image-to-text (I2T) understanding and text-to-image (T2I) generation are two fundamental, important yet traditionally isolated multimodal tasks. Despite their intrinsic connection, existing approaches typically optimize them independently, missing the opportunity for mutual enhancement. In this paper, we argue that the both tasks can be connected under a shared Auto-Encoder perspective, where text serves as the intermediate latent representation bridging the two directions — encoding images into textual semantics (I2T) and decoding text back into images (T2I). Our key insight is that if the encoder truly “understands” the image, it should capture all essential structure, and if the decoder truly “understands” the text, it should recover that structure fa

Unified Multimodal Models as Auto-Encoders Zhiyuan Yan 1,2 ⋄,⋆ , Kaiqing Lin 1 ⋄ , Zongjian Li 1,3 ⋄ , Junyan Ye 4 ⋄ , Hui Han 1 , Haochen Wang 2,6 ⋆ , Zhendong Wang 5 , Bin Lin 1,3 , Hao Li 1 , Xinyan Xiao 2 , Jingdong Wang 2 , Haifeng Wang 2 , Li Yuan 1 † 1 Shenzhen Graduate School, Peking University 2 Baidu, 3 Rabbitpre AI 4 SYSU, 5 USTC, 6 CASIA zhiyuanyan@stu.pku.edu.cn Abstract † † ⋄ \diamond Equal Contribution, ⋆ \star Work done during an internship at Baidu Star Program, † Corresponding Author Image-to-text (I2T) understanding and text-to-image (T2I) generation are two fundamental, imp

Explore this link on the map →

saved by

related reading