Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
Molmo and PixMo : Open Weights and Open Data for State-of-the-Art Vision-Language Models Matt Deitke ∗†ψ Christopher Clark ∗† Sangho Lee † Rohun Tripathi † Yue Yang † Jae Sung Park ψ Mohammadreza Salehi ψ Niklas Muennighoff † Kyle Lo † Luca Soldaini † Jiasen Lu † Taira Anderson † Erin Bransom † Kiana Ehsani † Huong Ngo † YenSung Chen † Ajay Patel † Mark Yatskar † Chris Callison-Burch † Andrew Head † Rose Hendrix † Favyen Bastani † Eli VanderBilt † Nathan Lambert † Yvonne Chou † Arnavi Chheda † Jenna Sparks † Sam Skjonsberg † Michael Schmitz † Aaron Sarnat † Byron Bischoff † Pete Walsh † Chris
Explore this link on the map →saved by
related reading
- 2403.09611.pdfarxiv.org
- Inkling: Our Open-Weights Model - Thinking Machines Labthinkingmachines.ai
- gpt-4.pdfcdn.openai.com
- GPT-4openai.com
- MolmoAct Action Reasoning Models that can Reason in Spacearxiv.org
- Seeing Is Not Reasoning: How VLMs and Their Benchmarks Lean on Textharvey-fin.github.io
- Composer2.pdfcursor.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Olmo from Ai2allenai.org
- Training VLM for CUA — Tzafontzafon.ai
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai