Seeing Is Not Reasoning: How VLMs and Their Benchmarks Lean on Text
We stress-test 8 open-weight VLMs on 9 multimodal benchmarks under different input conditions to ask how much each benchmark actually depends on the image, and propose per-instance metrics — Modality Differential, Caption Substitution, Visual Dependence — to audit it.
📌 TL;DR We test 8 openly released VLMs on 9 VQA benchmarks under different input conditions, and propose three metrics ( Modality Differential , Caption Substitution , Visual Dependence ) to ask two questions: how much does each benchmark actually depend on the image , and how much do VLMs actually reason over it? Many widely used VQA benchmarks barely need the image. In roughly a third of the (model × benchmark) pairs, the VLM performs as well or better when the image is replaced by a short caption. Visual dependence decreases as VLMs scale up, since a larger backbone answers more from
Explore this link on the map →saved by
related reading
- 2403.09611.pdfarxiv.org
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Modelsarxiv.org
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- GPT-4openai.com
- gpt-4.pdfcdn.openai.com
- Explore | alphaXivalphaxiv.org
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Part 1: The Map Was Wrong — Nemostationnemostation.com
- Inkling: Our Open-Weights Model - Thinking Machines Labthinkingmachines.ai
- [2209.15162] Linearly Mapping from Image to Text Spacearxiv.org
- cs.unc.edu/~mbansal/teaching/nlp-comp790-590-spring23.htmlcs.unc.edu
- VRPRM: Process Reward Modeling via Visual Reasoningarxiv.org