ResearchPod Summary
This paper investigates whether current vision-language models (VLMs) can correctly use spatial deictic expressions—words like "this" and "that"—which depend on the physical distance between the speaker and an object. While humans naturally adjust these expressions based on context and distance, it is unclear if VLMs possess this spatial-linguistic grounding. The authors developed a new benchmark based on the "memory game" paradigm, where models were asked to describe objects placed at varying distances (0.25m, 1.50m, and 2.75m) in four different languages: English, Japanese, Korean, and Chinese.
The experiments revealed that tested open-source VLMs (Gemma 3 and Qwen3-VL series) do not use demonstratives in a human-like manner. While human subjects consistently shift their choice of demonstrative as an object moves further away, the models showed a much weaker or non-existent correlation between distance and word choice. Furthermore, in languages with three demonstratives (Japanese and Korean), models exhibited a strong bias, often failing to use distal demonstratives even when appropriate. The study also noted that these models frequently struggled with the basic task of correctly identifying the shape and color of the objects, suggesting that their spatial reasoning is hindered by both visual recognition limitations and a lack of pragmatic understanding of language.
Spatial deixis is fundamental to human communication, bridging the gap between language and the physical environment. If VLMs are to be used in robotics or real-world navigation, they must be able to ground these expressions accurately. This research highlights a significant gap in the current capabilities of multimodal models, suggesting that their spatial reasoning is not yet sufficiently grounded in the physical reality of the scenes they process.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.