[2301.13823] Grounding Language Models to Images for Multimodal Generation
We propose an efficient method to ground pretrained text-only language models to the visual domain, enabling them to process and generate arbitrarily interleaved image-and-text data. Our method leverages the abilities of language models learnt from large scale text-only pretraining, such as in-context learning and free-form text generation. We keep the language model frozen, and finetune input and output linear layers to enable cross-modality interactions. This allows our model to process arbitrarily interleaved image-and-text inputs, and generate free-form text interleaved with retrieved images. We achieve strong zero-shot performance on grounded tasks such as contextual image retrieval and multimodal dialogue, and showcase compelling interactive abilities. Our approach works with any off-the-shelf language model and paves the way towards an effective, general solution for leveraging pretrained language models in visually grounded settings.
Grounding Language Models to Images for Multimodal Inputs and Outputs Jing Yu Koh 1 Ruslan Salakhutdinov 1 Daniel Fried 1 Abstract arXiv:2301.13823v4 [cs.CL] 13 Jun 2023 We propose an efficient method to ground pre- trained text-only language models to the visual domain, enabling them to process…
saved by
related reading
- MDETR - Modulated Detection for End-to-End Multi-Modal Understandingarxiv.org
- 2403.09611.pdfarxiv.org
- [2005.14165] Language Models are Few-Shot Learnersarxiv.org
- Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning | Researchai.meta.com
- Unified Multimodal Models as Auto-Encodersarxiv.org
- cs.unc.edu/~mbansal/teaching/nlp-comp790-590-spring23.htmlcs.unc.edu
- Minigpt-4minigpt-4.github.io
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversiontextual-inversion.github.io
- GitHub - FusionBrainLab/Vision_GRPOgithub.com
- Composing Zero-Shot Multimodal Reasoning with Languagesocraticmodels.github.io
- [2209.15162] Linearly Mapping from Image to Text Spacearxiv.org
- MMHal Bencharxiv.org