[2209.15162] Linearly Mapping from Image to Text Space
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2209.15162] Linearly Mapping from Image to Text Space Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Computation and Language arXiv:2209.15162 (cs) [Submitted on 30 Sep 2022 ( v1 ), last revised 9 Mar 2023 (this version, v3)] Title: Linearly Mapping from Image to Text Space Authors: Jack Merullo , Louis Castricato , Carsten Eickhoff , Ellie Pavlick View a PDF of the paper titled Linearly Mapping from Image to Text Space, by Jack Merullo and 3 other authors View PDF Abstract: The
Explore this link on the map →related reading
- [2301.13823] Grounding Language Models to Images for Multimodal Generationarxiv.org
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub
- [2103.00020] Learning Transferable Visual Models From Natural Language Supervisionarxiv.org
- 2403.09611.pdfarxiv.org
- Seeing Is Not Reasoning: How VLMs and Their Benchmarks Lean on Textharvey-fin.github.io
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Modelsarxiv.org
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversiontextual-inversion.github.io
- cs.unc.edu/~mbansal/teaching/nlp-comp790-590-spring23.htmlcs.unc.edu
- Turing Bletchley: A Universal Image Language Representation model by Microsoft - Microsoft Researchmicrosoft.com
- High-level visual representations in the human brain are aligned with large language models | Nature Machine Intelligencenature.com
- Multimodal interpretability in 2024soniajoseph.ai