flâneur — a map of the web's best reading

Seeing Is Not Reasoning: How VLMs and Their Benchmarks Lean on Text

harvey-fin.github.io · 3,704 words · saved by 1 readers

We stress-test 8 open-weight VLMs on 9 multimodal benchmarks under different input conditions to ask how much each benchmark actually depends on the image, and propose per-instance metrics — Modality Differential, Caption Substitution, Visual Dependence — to audit it.

📌 TL;DR We test 8 openly released VLMs on 9 VQA benchmarks under different input conditions, and propose three metrics ( Modality Differential , Caption Substitution , Visual Dependence ) to ask two questions: how much does each benchmark actually depend on the image , and how much do VLMs actually reason over it? Many widely used VQA benchmarks barely need the image. In roughly a third of the (model × benchmark) pairs, the VLM performs as well or better when the image is replaced by a short caption. Visual dependence decreases as VLMs scale up, since a larger backbone answers more from

Explore this link on the map →

saved by

related reading