CVJul 2

Domain Generalization via Text-Anchored Information Bottleneck

arXiv:2607.016576.7
Predicted impact top 64% in CV · last 90 daysOriginality Highly original
AI Analysis

For domain generalization in visual recognition, this work shifts the focus from improving representations to designing supervision that enforces invariance.

The paper identifies that preserving expressive visual representations from large vision-language models can propagate spurious cues that hinder domain generalization. By using language embeddings as an information bottleneck, they achieve state-of-the-art performance across diverse backbones.

Visual recognition models often fail when deployed in new environments. Domain Generalization (DG) addresses this by learning representations that remain invariant to environment-specific variations. Recent approaches increasingly rely on large vision-language models, assuming that preserving their expressive visual representations improves robustness. However, we show that such visual expressiveness can instead propagate spurious cues that tie representations to the training environments, hindering invariant learning. We therefore discard visual guidance and instead treat the language embedding space as the primary source of domain invariance, naturally acting as an information bottleneck that preserves core semantics while suppressing domain-specific variations. Extensive experiments across diverse backbones exhibit state-of-the-art performance and further analyze what makes guidance effective for robust generalization. These findings shift the focus of DG from improving representations to designing supervision that enforces invariance.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes