[2301.13823] Grounding Language Models to Images for Multimodal Generation
We propose an efficient method to ground pretrained text-only language models to the visual domain, enabling them to process and generate arbitrarily interleaved image-and-text data. Our method leverages the abilities of language models learnt from large scale text-only pretraining, such as in-context learning and free-form text generation. We keep the language model frozen, and finetune input and output linear layers to enable cross-modality interactions. This allows our model to process arbitrarily interleaved image-and-text inputs, and generate free-form text interleaved with retrieved images. We achieve strong zero-shot performance on grounded tasks such as contextual image retrieval and multimodal dialogue, and showcase compelling interactive abilities. Our approach works with any off-the-shelf language model and paves the way towards an effective, general solution for leveraging pretrained language models in visually grounded settings.
We propose an efficient method to ground pretrained text-only language models to the visual domain, enabling them to process and generate arbitrarily interleaved image-and-text data. Our method leverages the abilities of language models learnt from large scale text-only pretraining, such as in-context learning and free-form text generation. We keep the language model frozen, and finetune input and output linear layers to enable cross-modality interactions. This allows our model to process arbitrarily interleaved image-and-text inputs, and generate free-form text interleaved with retrieved imag
Explore this link on the map →saved by
related reading
- MDETR - Modulated Detection for End-to-End Multi-Modal Understandingarxiv.org
- Multi-View Transformer for 3D Visual Groundingarxiv.org
- 2403.09611.pdfarxiv.org
- Unified Multimodal Models as Auto-Encodersarxiv.org
- cs.unc.edu/~mbansal/teaching/nlp-comp790-590-spring23.htmlcs.unc.edu
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversiontextual-inversion.github.io
- Replicate - Run AI with an APIreplicate.com
- Explore | alphaXivalphaxiv.org
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Modelsarxiv.org
- [2209.15162] Linearly Mapping from Image to Text Spacearxiv.org
- The Illustrated Stable Diffusion – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Introducing CM3leon, a more efficient, state-of-the-art generative model for text and imagesai.meta.com