CVNov 13, 2021

Factorial Convolution Neural Networks

arXiv:2111.07072v1

Originality Incremental advance

AI Analysis

This work addresses efficiency and accuracy challenges in object detection for resource-starved applications, representing an incremental improvement over existing methods.

The paper tackles the issues of contaminated deep features and high execution overheads in GoogleNet for object detection by proposing FactorNet, a CNN composed of multiple independent sub-CNNs that encode different aspects of visual features. Incorporating FactorNet into Faster-RCNN resulted in at least a 5% better accuracy and additional speedup over GoogleNet on the KITTI benchmark dataset.

In recent years, GoogleNet has garnered substantial attention as one of the base convolutional neural networks (CNNs) to extract visual features for object detection. However, it experiences challenges of contaminated deep features when concatenating elements with different properties. Also, since GoogleNet is not an entirely lightweight CNN, it still has many execution overheads to apply to a resource-starved application domain. Therefore, a new CNNs, FactorNet, has been proposed to overcome these functional challenges. The FactorNet CNN is composed of multiple independent sub CNNs to encode different aspects of the deep visual features and has far fewer execution overheads in terms of weight parameters and floating-point operations. Incorporating FactorNet into the Faster-RCNN framework proved that FactorNet gives \ignore{a 5\%} better accuracy at a minimum and produces additional speedup over GoolgleNet throughout the KITTI object detection benchmark data set in a real-time object detection system.

View on arXiv PDF

Similar