CVJul 9

Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation

arXiv:2607.083978.9h-index: 10
Predicted impact top 47% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the lack of high-quality annotations and domain-specific challenges in endoscopic referring segmentation for the medical imaging community.

The authors introduce ReferEndoscopy, a large-scale benchmark for referring image segmentation in endoscopy, and propose AR-ERIS, an attribute retrieval-based framework that achieves state-of-the-art performance with strong generalization across simulated and real-world endoscopic data.

Referring Image Segmentation (RIS) aims to segment image regions specified by natural language, enabling fine-grained and controllable visual understanding. Extending RIS to endoscopic imagery, however, presents unique challenges, including scarce high-quality annotations and complex, domain-specific image-text relationships. Although recent vision-language models demonstrate strong cross-domain alignment, they often fail to capture fine-grained textual cues in endoscopic settings, resulting in suboptimal performance and limited generalization. To address these challenges, we introduce ReferEndoscopy, a large-scale benchmark for RIS in the endoscopy field. Building on this dataset, we propose the Attribute Retrieval-based Endoscopic-RIS (AR-ERIS) framework for open-vocabulary endoscopic compositional referring segmentation. AR-ERIS leverages attribute retrieval for open-vocabulary endoscopic compositional referring segmentation and is pretrained on the curated ReferEndoscopy dataset, achieving state-of-the-art performance with strong generalization across both simulated and real-world endoscopic data. The dataset and code will be publicly released upon completion of the review process.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes