flâneur

Circuit Tracing in Vision–Language Models: Understanding the Internal Mechanisms of Multimodal Thinking

arxiv.org · 8,057 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. Vision–language models (VLMs) are powerful but remain opaque black boxes. We introduce the first framework for transparent circuit tracing in VLMs to systematically analyze multimodal reasoning. By utilizing transcoders, attribution graphs, and attention-based methods, we uncover how VLMs hierarchically integrate visual and semantic concepts. We reveal that distinct visual feature circuits can handle mathematical reasoning and support cross-modal associations. Validated through feature steering and circuit patching, our framework proves these circuits are causal and controllable, laying the groundwork for more explainable and reliable VLMs. Our code and models are available at https://github.com/UIUC-MONET/vlm-circuit-tracing. The rapid advancement of vi

Tianhu Xiong Email: mw34@illinois.edu Shengyi Qian Email: klara@illinois.edu Klara Nahrstedt Affiliation: University of Illinois Urbana-Champaign Mingyuan Wu Affiliation: University of Illinois Urbana-Champaign Affiliation: Independent Researcher Abstract Vision–language models (VLMs) are powerful but remain opaque black boxes. We introduce the first framework for transparent circuit tracing in VLMs to systematically analyze multimodal reasoning. By utilizing transcoders, attribution graphs, and attention-based methods, we uncover how VLMs hierarchically integrate visual and…

saved by

related reading