Turing Bletchley: A Universal Image Language Representation model by Microsoft - Microsoft Research
Today, the Microsoft Turing team is thrilled to introduce Turing Bletchley, a 2.5-billion parameter Universal Image Language Representation model (T-UILR) that can perform image-language tasks in 94 languages. T-Bletchley has an image encoder and a universal language encoder that vectorize input image and text respectively so that semantically similar images and texts align with each other. This model shows […]
Turing Bletchley: A Universal Image Language Representation model by Microsoft - Microsoft Research Skip to main content Research Publications Code & data People Microsoft Research blog Artificial intelligence Audio & acoustics Computer vision Graphics & multimedia Human-computer interaction Human language technologies Search & information retrieval Data platforms and analytics Hardware & devices Programming languages & software engineering Quantum computing Security, privacy & cryptography Systems & networking Algorithms Mathematics Ecology & environment Economics Medical, health & genomics S
Explore this link on the map →related reading
- Inkling: Our Open-Weights Model - Thinking Machines Labthinkingmachines.ai
- 2403.09611.pdfarxiv.org
- [2301.13823] Grounding Language Models to Images for Multimodal Generationarxiv.org
- Unified Multimodal Models as Auto-Encodersarxiv.org
- Replicate - Run AI with an APIreplicate.com
- [2209.15162] Linearly Mapping from Image to Text Spacearxiv.org
- cs.unc.edu/~mbansal/teaching/nlp-comp790-590-spring23.htmlcs.unc.edu
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversiontextual-inversion.github.io
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Modelsarxiv.org
- Generalized Language Models | Lil'Loglilianweng.github.io
- Training VLM for CUA — Tzafontzafon.ai
- Fast and Simple Image Search with Foundation Models - Ivan Zhouivanzhou.me