CVJun 4

Anchored, Not Graded: Vision-Language Models Fail at Slant-from-Texture Perception

arXiv:2606.067148.8
Predicted impact top 10% in CV · last 90 daysOriginality Incremental advance
AI Analysis

Identifies a fundamental limitation in VLMs' ability to express low-level geometric cues in a graded manner, relevant for understanding their perceptual competences.

Vision-Language Models (VLMs) fail to reproduce human-like graded slant-from-texture perception, instead predicting slant only at discrete anchor values (e.g., 0°, ±25°, ±45°) with little sensitivity to stimulus variations. Supervised fine-tuning only partially remediates this anchoring failure.

Human perception of surface slant from texture exhibits systematic, graded biases that emerge reliably in psychophysical experiments. Prior work showed that unsupervised CNNs reproduce several human-like biases, while supervised CNNs do not. Do Vision-Language Models (VLMs) exhibit similar competences? Across multiple VLM families and model scales, zero-shot and in-context prompting both produce distinctive failures: slant is predicted at only a small set of anchors (e.g., 0\degree, $\pm$25\degree, $\pm$45\degree) with little dependence on stimulus field of view, optical slant, or surface curvature. Supervised fine-tuning partially remediates the failure, but residual anchoring persists. While success in high-level vision-language benchmarks might not require sensitivity to low-level geometric cues, we interpret anchoring as a failure at the representation-to-output language interface: Not necessarily an absence of geometric encoding, but a failure to express it in a graded form.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes