Rethinking Feature Reliance Evaluation with Semantically Matched Suppression
For researchers studying model interpretability and robustness, this work provides a more rigorous method for evaluating feature reliance, revealing that prior conclusions about CNN texture bias may be artifacts of unmatched suppression.
The paper introduces a semantically matched evaluation framework for comparing shape and texture suppression in visual recognition models, showing that ImageNet-trained CNNs exhibit greater texture reliance than previously thought, and that Vision Transformers retain higher accuracy under both types of suppression.
Understanding whether visual recognition models rely on shape, texture, or color is central to interpreting their behavior. Prior cue-conflict studies have strongly influenced the view that CNNs are texture-biased, yet such tests measure cue preference under artificial conflicts rather than feature reliance during natural recognition. We revisit this question through controlled feature suppression and show that performance drops are difficult to interpret unless different suppression operations impose comparable category-level damage. We introduce a semantically matched evaluation framework that compares shape and texture suppression at matched levels of category separability loss. Under this framework, ImageNet-trained CNNs show stronger degradation under texture suppression than under shape suppression, revealing greater texture reliance than suggested by unmatched suppression analyses. Extending the comparison across architectures, we find that Vision Transformers retain higher accuracy than CNNs under both shape and texture suppression. Brain encoding further shows that ViT representations exhibit smaller suppression-induced decreases in neural prediction performance under the tested suppression settings. These findings indicate that semantic comparability is essential for interpreting feature reliance from suppression experiments, and suggest that the robustness advantage of ViTs may be related to representations more compatible with human visual cortex.