✳flâneur — a map of the web's best reading
2403.09611.pdf
arxiv.org · 17,221 words · saved by 2 readers
N/A
# link_f61mfa7zvt.pdf ## Metadata - PDFFormatVersion=1.6 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - CreationDate=D:20240422000405Z - Creator=LaTeX with hyperref - ModDate=D:20240422000405Z - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.25 (TeX Live 2023) kpathsea version 6.3.5 - Producer=pdfTeX-1.40.25 - Trapped=False ## Contents ### Page 1 MM1: Methods, Analysis & Insights from Multimodal LLM Pre-trainingBrandon McKinzie◦, Zhe Gan◦, Jean-Philippe Fauconnier⋆, Sam Dodge⋆, Bowen Zhang⋆,
Explore this link on the map →saved by
related reading
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Modelsarxiv.org
- Inkling: Our Open-Weights Model - Thinking Machines Labthinkingmachines.ai
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- [2301.13823] Grounding Language Models to Images for Multimodal Generationarxiv.org
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Large Language Diffusion Modelsarxiv.org
- Seeing Is Not Reasoning: How VLMs and Their Benchmarks Lean on Textharvey-fin.github.io
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrouai.googleblog.com
- Training Multimodalnimapourjafar.com
- The Llama Hitchiking Guide to Local LLMs – hackerllamaosanseviero.github.io