CVJul 1

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning

arXiv:2607.0119127.1
Predicted impact top 2% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For vision-language models tackling fine-grained visual reasoning, this work introduces a principled decoupling of perception and reasoning with a role-aware RL strategy, showing consistent gains across scales and benchmarks.

The paper proposes Perceive-to-Reason (P2R), a two-stage framework that decouples perception and reasoning for fine-grained visual reasoning, achieving 93.2% on V-Star, 81.9% on HR-Bench-4K, and 80.5% on HR-Bench-8K with a 4B model, outperforming its backbone.

Fine-grained visual reasoning remains challenging for vision-language models, especially when small but critical visual cues are buried in high-resolution images. Existing approaches rely on repeated cropping or test-time visual search to introduce local evidence, but they typically do not explicitly distinguish perception from reasoning. In this paper, we propose Perceive-to-Reason (P2R), a unified framework that formulates fine-grained visual reasoning as a two-stage process: the model first localizes question-relevant evidence as a Perceiver, and then answers the question as a Reasoner based on the annotated image and cropped regions. To better align training with this decoupled formulation, we further introduce Perception-Reasoning Alternating GRPO (PRA-GRPO), a role-aware reinforcement learning strategy that alternates between perception-focused and reasoning-focused updates using only final-answer supervision. Built on top of Qwen3-VL-Instruct-2B/4B/8B, P2R consistently improves performance across model scales. In particular, P2R-4B achieves 93.2% on V-Star, 81.9% on HR-Bench-4K, and 80.5% on HR-Bench-8K, substantially outperforming its corresponding backbone. Further experiments show that the benefits of P2R extend beyond high-resolution benchmarks to broader multimodal reasoning tasks. These results suggest that explicitly decoupling perception from reasoning provides an effective framework for fine-grained visual reasoning.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes