Seeing Is Not Reasoning: How VLMs and Their Benchmarks Lean on Text
We stress-test 8 open-weight VLMs on 9 multimodal benchmarks under different input conditions to ask how much each benchmark actually depends on the image, and propose per-instance metrics — Modality Differential, Caption Substitution, Visual Dependence — to audit it.
📌 TL;DR We test 8 openly released VLMs on 9 VQA benchmarks under different input conditions, and propose three metrics ( Modality Differential , Caption Substitution , Visual Dependence ) to ask two questions: how much does each benchmark actually depend on the image , and how much do VLMs actually reason over it? Many widely used VQA benchmarks barely need the image. In roughly a third of the (model × benchmark) pairs, the VLM performs as well or better when the image is replaced by a short caption. Visual dependence decreases as VLMs scale up, since a larger backbone answers more from
saved by
related reading
- 2403.09611.pdfarxiv.org
- Less Detail, Better Answers: Degradation-Driven Prompting for VQAarxiv.org
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Modelsarxiv.org
- MMHal Bencharxiv.org
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- Circuit Tracing in Vision–Language Models:Understanding the Internal Mechanisms of Multimodal Thinkingarxiv.org
- Explore | alphaXivalphaxiv.org
- TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answeringarxiv.org
- [2301.13823] Grounding Language Models to Images for Multimodal Generationarxiv.org
- Part 1: The Map Was Wrong — Nemostationnemostation.com
- cs.unc.edu/~mbansal/teaching/nlp-comp790-590-spring23.htmlcs.unc.edu
- [2209.15162] Linearly Mapping from Image to Text Spacearxiv.org