CVJul 9

LOGOS: Language-guided Oriented Object Detection in Aerial Scenes

arXiv:2607.080043.8h-index: 3
Predicted impact top 84% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For remote sensing applications, LOGOS improves oriented object detection accuracy in complex scenes, but the gains are incremental over existing methods.

LOGOS introduces a transformer-based method using textual prompts to guide oriented object detection in aerial scenes, outperforming state-of-the-art on the DOTA dataset, especially for densely packed and rotated objects.

Object detection in geospatial scenes, such as satellite and aerial imagery, poses significant challenges due to the varying orientations and densities of objects, as well as the complex backgrounds inherent to remote sensing imagery. Traditional methods for oriented object detection have struggled to address issues such as angular discontinuity, fixed query sizes, and inefficiencies in handling sparse or cluttered scenes. In this paper, we propose LOGOS, a novel transformer-based approach that leverages textual prompts to guide the detection of oriented objects in aerial scenes. In particular, our proposed approach incorporates prompt-modulated content queries to dynamically adjust the model's focus based on the provided text, thereby improving object detection accuracy in complex environments. Empirically, extensive experiments on the DOTA dataset demonstrate that LOGOS outperforms existing state-of-the-art methods, particularly in densely packed and rotated object scenarios. Our approach offers a significant step forward in improving the robustness and scalability of oriented object detection in remote sensing applications.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes