CLJul 8

Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models

arXiv:2607.0725116.8h-index: 12
Predicted impact top 43% in CL · last 90 daysOriginality Synthesis-oriented
AI Analysis

For researchers evaluating spatial reasoning in VLMs, this work provides a new benchmark but is incremental as it applies existing evaluation concepts to a specific linguistic phenomenon.

The paper develops a benchmark to evaluate vision-language models' ability to use spatial deictic expressions (e.g., 'this', 'that') in four languages, finding that models differ from humans in selecting demonstratives based on object distance.

One of the expected abilities of vision-language models (VLMs) is spatial reasoning ability based on a given text and image. To evaluate the spatial reasoning abilities of VLMs, we focus on the use of spatial deictic expressions, which are defined as spatial expressions whose referent is determined by their situational context, such as ``this'' and ``that''. To handle spatial deictic expressions, VLMs must jointly reason over language and visual space, grounding context-dependent references in the image's spatial structure. In addition, selecting appropriate spatial deictic expressions across languages requires VLMs to understand the language-specific spatial distinctions encoded by these expressions. In this paper, we develop a benchmark to evaluate the multilingual ability of VLMs to use spatial deictic expressions in four languages. Our experiments using this benchmark reveal that the tested models use demonstratives in a manner different from that of humans, particularly in selecting the appropriate demonstratives based on the distance to the object.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes