Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning | Meta AI Research
ai.meta.com · 290 words · saved by 1 readers
We present CM3Leon (pronounced “Chameleon”), a retrieval-augmented, tokenbased, decoder-only multi-modal language model capable of generating and...
NLP COMPUTER VISION July 14, 2023 Abstract We present CM3Leon (pronounced “Chameleon”), a retrieval-augmented, tokenbased, decoder-only multi-modal language model capable of generating and infilling both text and images. CM3Leon uses the CM3 multi-modal architecture but additionally shows the extreme benefits of scaling up and tuning on more diverse instruction-style data. It is the first multi-modal model trained with a recipe adapted from text-only language models, including a large-scale retrieval-augmented pretraining stage and a second multi-task supervised fine-tuning (SFT) stage.…
saved by
related reading
- Introducing CM3leon, a more efficient, state-of-the-art generative model for text and imagesai.meta.com
- 2403.09611.pdfarxiv.org
- [2301.13823] Grounding Language Models to Images for Multimodal Generationarxiv.org
- Unified Multimodal Models as Auto-Encodersarxiv.org
- [2005.14165] Language Models are Few-Shot Learnersarxiv.org
- Large Language Diffusion Modelsarxiv.org
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversiontextual-inversion.github.io
- 2408.11039arxiv.org
- MMHal Bencharxiv.org
- Chunyuan Lichunyuan.li
- Minigpt-4minigpt-4.github.io
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Modelsarxiv.org