CVAug 25, 2019

A Comparison of CNN and Classic Features for Image Retrieval

Umut Özaydın, Theodoros Georgiou, Michael Lew

arXiv:1908.09300v114 citations

AI Analysis

This work provides an incremental comparison for computer vision researchers and practitioners in image retrieval.

The paper compared CNN-based and conventional keypoint detection methods for image retrieval, finding that each type of features performs best in different contexts.

Feature detectors and descriptors have been successfully used for various computer vision tasks, such as video object tracking and content-based image retrieval. Many methods use image gradients in different stages of the detection-description pipeline to describe local image structures. Recently, some, or all, of these stages have been replaced by convolutional neural networks (CNNs), in order to increase their performance. A detector is defined as a selection problem, which makes it more challenging to implement as a CNN. They are therefore generally defined as regressors, converting input images to score maps and keypoints can be selected with non-maximum suppression. This paper discusses and compares several recent methods that use CNNs for keypoint detection. Experiments are performed both on the CNN based approaches, as well as a selection of conventional methods. In addition to qualitative measures defined on keypoints and descriptors, the bag-of-words (BoW) model is used to implement an image retrieval application, in order to determine how the methods perform in practice. The results show that each type of features are best in different contexts.

View on arXiv PDF

Similar