ROJun 17

ReSiReg: Towards Spatially Consistent Semantics in Language-Conditioned Robotic Tasks

arXiv:2606.1908810.9
Predicted impact top 35% in RO · last 90 daysOriginality Incremental advance
AI Analysis

For robotic manipulation requiring simultaneous semantic and 3D spatial reasoning, ReSiReg addresses the noise and lack of spatial consistency in VLM embeddings, enabling more reliable open-language instruction following.

ReSiReg improves spatial consistency of dense VLM embeddings for language-conditioned robotic tasks by reconstructing patch features as soft mixtures of prototype-level language descriptors, achieving better dense retrieval and spatially consistent activations in manipulation scenes, with a compact 25M model competitive with ViT-B baselines.

Vision-Language Models (VLMs) enable robots to follow open-language instructions. However, dense VLM embeddings have shown to be noisy and lack spatial consistency. This is problematic for robotic applications, which require simultaneous reasoning over semantics and 3D space. We examine spatial structure across recent VLMs and propose ReSiReg, a feature reconstruction method that uses spatially consistent VLM intermediates to improve dense language-grounded retrieval. ReSiReg clusters intermediates into visual prototypes, derives their language descriptors, and reconstructs each patch as a soft mixture of prototype-level language embeddings. We evaluate quantitatively on OVSS and 3D mapping across backbones, and qualitatively in real-world manipulation scenes. Quantitative results show improved dense retrieval; manipulation scenes show more spatially consistent target activations. We further provide a compact 25M dense VLM for robotic applications, substantially smaller than and competitive with ViT-B baselines. Available at https://resireg.github.io

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes