CVAIJun 25

Bridging Vision and Language Concepts through Optimal Transport Semantic Flow

arXiv:2606.2689111.0
Predicted impact top 43% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For researchers in interpretable AI and vision-language reasoning, this work introduces a new geometric perspective for concept alignment, though it is incremental as it builds on existing CBM frameworks.

The paper tackles the problem of aligning visual and textual representations in Concept Bottleneck Models (CBMs). It proposes OTF-CBM, which uses optimal transport flow matching to model semantic transitions, achieving superior classification accuracy and concept faithfulness.

Concept Bottleneck Models (CBMs) promise transparent reasoning by predicting through human-interpretable concepts, yet their effectiveness fundamentally depends on how well visual and textual representations are aligned or matched. Existing vision-language CBMs often rely on pre-aligned encoders or global cosine similarity, which obscures fine-grained concept localization and fails to reflect true semantic geometry. In this work, we rethink concept alignment as a dynamic cross-modal transport process instead of static projection and propose the Optimal Transport Flow Concept Bottleneck Model (OTF-CBM). It first learns a data-driven semantic cost via Inverse Optimal Transport to measure cross-modal distances, and then performs unbalanced optimal-transport-based flow matching to model semantic transitions between visual patches and textual concepts. With velocity-based concept activation, OTF-CBM captures interpretable geometric relations without ODE integration. Experiments further show that OTF-CBM achieves superior classification accuracy and concept faithfulness, offering a new geometric and dynamical perspective for interpretable cross-modal reasoning.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes