CVCLLGJul 3

Pathways of Visual Information Flow in Vision-Language Models

arXiv:2607.0335815.2
Predicted impact top 20% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For researchers studying mechanistic interpretability of multimodal models, this work provides a causal understanding of how visual information is routed, revealing flexibility that unifies prior findings.

The paper identifies two distinct pathways (direct and text-mediated) for visual information flow in vision-language models, showing that pathway selection is task-dependent and flexible under intervention.

We study how visual information is routed in vision-language models (VLMs). Using causal patching on controlled synthetic and natural datasets, we find that models rely on two distinct pathways to solve visual tasks: A direct pathway, where visual information is retained in image token representations and read out by the final token at later layers, and a text-mediated pathway, where visual information is first transferred to the query tokens and then read out by the final token. Across three visual tasks, we show that pathway selection is task-dependent, and that data distribution and prompt design can also modulate which pathway is used to solve the image-based query. Moreover, using attention knockouts and corrupted-input patching, we find that these pathways are flexible, under certain interventions, models can rely on the text-mediated pathway as a fallback when the usual pathway is ablated. This behavior unifies findings in prior work and shows that ablation-based interventions can reveal what models could do rather than what they normally do. Together, our results provide a mechanistic characterization of visual information flow in VLMs and highlight the flexibility of their internal mechanisms under intervention.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes