CV AIMay 2, 2025

CDFormer: Cross-Domain Few-Shot Object Detection Transformer Against Feature Confusion

Boyuan Meng, Xiaohan Zhang, Peilin Li, Zhe Wu, Yiming Li, Wenkai Zhao, Beinan Yu, Hui-Liang Shen

arXiv:2505.00938v111.84 citationsh-index: 7Has CodeICME

Originality Incremental advance

AI Analysis

It solves object detection across domains with limited data, which is incremental as it builds on existing methods with specific modules.

The paper tackled cross-domain few-shot object detection by addressing feature confusion, resulting in CDFormer achieving improvements of 12.9%, 11.0%, and 10.4% mAP under 1/5/10 shot settings.

Cross-domain few-shot object detection (CD-FSOD) aims to detect novel objects across different domains with limited class instances. Feature confusion, including object-background confusion and object-object confusion, presents significant challenges in both cross-domain and few-shot settings. In this work, we introduce CDFormer, a cross-domain few-shot object detection transformer against feature confusion, to address these challenges. The method specifically tackles feature confusion through two key modules: object-background distinguishing (OBD) and object-object distinguishing (OOD). The OBD module leverages a learnable background token to differentiate between objects and background, while the OOD module enhances the distinction between objects of different classes. Experimental results demonstrate that CDFormer outperforms previous state-of-the-art approaches, achieving 12.9% mAP, 11.0% mAP, and 10.4% mAP improvements under the 1/5/10 shot settings, respectively, when fine-tuned.

View on arXiv PDF Code

Similar