CVJul 1

GEAR-Seg: A Grounded Explainable Agent for Reasoning Segmentation and Data Engine

arXiv:2607.0054410.4
Predicted impact top 38% in CV · last 90 daysOriginality Highly original
AI Analysis

For the reasoning segmentation community, GEAR-Seg provides an interpretable and scalable alternative to opaque end-to-end models, with a data engine that reduces reliance on costly human annotations.

GEAR-Seg decouples reasoning segmentation into class-agnostic segmentation, semantic description, and LLM deduction, achieving competitive zero-shot performance across multiple benchmarks. It also serves as a data engine to generate GEAR-131K, a large-scale benchmark, and enables lightweight models to approach human-annotated baseline performance.

Reasoning segmentation requires localizing targets based on complex, implicit queries. Current end-to-end models typically entangle perception and deduction into an opaque black box, severely limiting interpretability and scalability. To address this, we propose GEAR-Seg (Grounded Explainable Agent for Reasoning Segmentation), an explicitly decoupled agent that shifts the paradigm by translating visual pixels into dense, attribute-rich text. By decoupling class-agnostic segmentation, semantic description, and Large Language Model (LLM) deduction, GEAR-Seg transforms implicit reasoning into an explicit, trackable logic chain. As a zero-shot inference framework, it achieves highly competitive performance across diverse reasoning and fine-grained referring segmentation benchmarks. Furthermore, GEAR-Seg inherently functions as a highly scalable data engine. Utilizing this engine, we construct GEAR-131K, a massive benchmark (over 38k images, 656k QA-mask pairs) introducing a multifaceted taxonomy tailored for complex real-world manipulation-oriented reasoning. Finally, distillation experiments demonstrate that lightweight models supervised exclusively by our automated pipeline closely match the upper-bound performance of costly human-annotated baselines.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes