Training Multimodal
TLDR: We achieve competitive results in multimodal AI using late fusion techniques, training an 8B parameter model on a single 8xH100 cluster, outperforming similar-sized open-source models across various benchmarks for less than $2000 We also test out an early fusion architecture and determine that while it initially underperforms compared to late fusion, it shows potential for richer cross-modal learning given more extensive training Want to just see our experimental code: check it out here! Want to start training your own model? Check out the Omega Labs subnet! Get paid (enough to cover compute) to train your own models I also have a bunch of multimodal datasets you can check out on Hugging Face exported to conversation formats I spent a good chunk of summer at Omega Labs helping them kickoff their experiments with multimodal artificial intelligence — from creating datasets, setting up training code, and doing the runs A big issue I found while doing this is the lack of resources an
Training Multimodal > Nima Pourjafar [ projects ] [ blog ] TLDR: We achieve competitive results in multimodal AI using late fusion techniques, training an 8B parameter model on a single 8xH100 cluster, outperforming similar-sized open-source models across various benchmarks for less than $2000 We also test out an early fusion architecture and determine that while it initially underperforms compared to late fusion, it shows potential for richer cross-modal learning given more extensive training Want to just see our experimental code: check it out here! Want to start training your own model? Che
Explore this link on the map →related reading
- 2403.09611.pdfarxiv.org
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- ImageBind: Holistic AI learning across six modalitiesai.facebook.com
- Interaction Models: A Scalable Approach to Human-AI Collaboration - Thinking Machines Labthinkingmachines.ai
- MAS.S60 How2AI | Schedulemit-mi.github.io
- Unified Multimodal Models as Auto-Encodersarxiv.org
- Composer2.pdfcursor.com
- Training great LLMs entirely from ground up in the wilderness as a startup - Yi Tayyitay.net
- [2606.02800] Cosmos 3: Omnimodal World Models for Physical AIarxiv.org
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- MDETR - Modulated Detection for End-to-End Multi-Modal Understandingarxiv.org
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Modelsarxiv.org