flâneur

Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning | Meta AI Research

ai.meta.com · 290 words · saved by 1 readers

We present CM3Leon (pronounced “Chameleon”), a retrieval-augmented, tokenbased, decoder-only multi-modal language model capable of generating and...

NLP COMPUTER VISION July 14, 2023 Abstract We present CM3Leon (pronounced “Chameleon”), a retrieval-augmented, tokenbased, decoder-only multi-modal language model capable of generating and infilling both text and images. CM3Leon uses the CM3 multi-modal architecture but additionally shows the extreme benefits of scaling up and tuning on more diverse instruction-style data. It is the first multi-modal model trained with a recipe adapted from text-only language models, including a large-scale retrieval-augmented pretraining stage and a second multi-task supervised fine-tuning (SFT) stage.…

saved by

related reading