[2606.02800] Cosmos 3: Omnimodal World Models for Physical AI
Abstract:We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture. By supporting highly flexible input-output configurations, Cosmos 3 seamlessly unifies critical modalities for Physical AI -- effectively subsuming vision-language models, video generators, world simulators, and world-action models into a single framework. Our evaluation demonstrates that Cosmos 3 establishes a new state-of-the-art across a diverse suite of understanding and generation tasks, demonstrating omnimodal world models as scalable, general-purpose backbones for embodied agents. Our post-trained Cosmos 3 models were ranked as the best open-source Text-to-Image and Image-to-Video models by Artificial Analysis, and the best policy model by RoboArena at the time the technical report was written. To accelerate open research and deployment in Physical AI, we make our code, model checkpoints, curated synthetic datasets, and evaluation benchmark available under the Linux Foundation's OpenMDW-1.1 License at this https URL and this https URL. The project website is available at this https URL.
# link_1ag7y2ccllr.pdf ## Metadata - PDFFormatVersion=1.7 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=NVIDIA; :; Aditi; Niket Agarwal; Arslan Ali; Jon Allen; Martin Antolini; Adeline Aubame; Alisson Azzolini; Junjie Bai; Maciej Bala; Yogesh Balaji; Josh Bapst; Aarti Basant; Mukesh Beladiya; Mohammad Qazim Bhat; Zaid Pervaiz Bhat; Dan Blick; Vanni Brighella; Han Cai; Tiffany Cai; Eric Cameracci; Jiaxin Cao; Yulong Cao; Mark Carlson; Carlos Casanova; Ting-Yun Chang; Yan Chang; Yu-Wei Chao; Prithvijit Chattop
Explore this link on the map →saved by
related reading
- A VLA with Open-World Generalizationpi.website
- World Models | Rohit Bandarurohitbandaru.github.io
- A Functional Taxonomy of World Models - Dr. Fei-Fei Lidrfeifei.substack.com
- Explore | alphaXivalphaxiv.org
- World Models: Computing the Uncomputablenotboring.co
- The flavor of the bitter lesson for computer vision - Vincent Sitzmannvincentsitzmann.com
- The Model That Dreams the Worldmoe-capital.com
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixelsle-wm.github.io
- Marble: A Multimodal World Model | World Labsworldlabs.ai
- Frontier Systems for the Physical World - by Oliver Hsua16z.news
- Unified Multimodal Models as Auto-Encodersarxiv.org