CVSCJun 21

SATURN: Symbolic Spatial Reasoning for Multi-Perspective Grounding

arXiv:2606.2269421.4
Predicted impact top 11% in CV · last 90 daysOriginality Highly original
AI Analysis

For vision-language models, SATURN addresses the bottleneck of unreliable spatial reasoning under varying frames of reference.

SATURN introduces a neuro-symbolic framework for perspective-aware compositional spatial reasoning that remains stable under increasing complexity, outperforming strong baselines by 14 percentage points on the real-world MindCube benchmark.

Vision-Language Models (VLMs) remain unreliable when spatial reasoning requires composing relations whose meanings depend on frames of reference. Existing neuro-symbolic methods make reasoning more explicit, but often depend on brittle geometric procedures and hard decisions over noisy perception. We propose SATURN, a neuro-symbolic framework for perspective-aware compositional spatial reasoning. SATURN reconstructs an approximate 3D scene, derives soft perspective-aware spatial predicates, and composes them with a training-free Pythonic symbolic executor, separating perception from reasoning while preserving uncertainty through multi-hop inference. We also introduce 3D FORCE, a diagnostic benchmark that controls reasoning depth, view, and perspective composition across spatial arrangement grounding (SAG) and referring expression grounding (REF). On 3D FORCE, VLMs and spatially trained models degrade sharply as depth and perspective complexity increase, whereas SATURN remains stable and outperforms strong baselines. On the real-world MindCube benchmark, SATURN achieves 78.57% overall accuracy, outperforming the strongest baseline by 14 pp.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes