Li Zhu

CV
h-index9
4papers
124citations
Novelty51%
AI Score34

4 Papers

13.0CVDec 11, 2019Code
IoU-uniform R-CNN: Breaking Through the Limitations of RPN

Li Zhu, Zihao Xie, Liman Liu et al.

Region Proposal Network (RPN) is the cornerstone of two-stage object detectors, it generates a sparse set of object proposals and alleviates the extrem foregroundbackground class imbalance problem during training. However, we find that the potential of the detector has not been fully exploited due to the IoU distribution imbalance and inadequate quantity of the training samples generated by RPN. With the increasing intersection over union (IoU), the exponentially smaller numbers of positive samples would lead to the distribution skewed towards lower IoUs, which hinders the optimization of detector at high IoU levels. In this paper, to break through the limitations of RPN, we propose IoU-Uniform R-CNN, a simple but effective method that directly generates training samples with uniform IoU distribution for the regression branch as well as the IoU prediction branch. Besides, we improve the performance of IoU prediction branch by eliminating the feature offsets of RoIs at inference, which helps the NMS procedure by preserving accurately localized bounding box. Extensive experiments on the PASCAL VOC and MS COCO dataset show the effectiveness of our method, as well as its compatibility and adaptivity to many object detection architectures. The code is made publicly available at https://github.com/zl1994/IoU-Uniform-R-CNN,

3.7CVDec 20, 2024
Sparse Point Clouds Assisted Learned Image Compression

Yiheng Jiang, Haotian Zhang, Li Li et al.

In the field of autonomous driving, a variety of sensor data types exist, each representing different modalities of the same scene. Therefore, it is feasible to utilize data from other sensors to facilitate image compression. However, few techniques have explored the potential benefits of utilizing inter-modality correlations to enhance the image compression performance. In this paper, motivated by the recent success of learned image compression, we propose a new framework that uses sparse point clouds to assist in learned image compression in the autonomous driving scenario. We first project the 3D sparse point cloud onto a 2D plane, resulting in a sparse depth map. Utilizing this depth map, we proceed to predict camera images. Subsequently, we use these predicted images to extract multi-scale structural features. These features are then incorporated into learned image compression pipeline as additional information to improve the compression performance. Our proposed framework is compatible with various mainstream learned image compression models, and we validate our approach using different existing image compression methods. The experimental results show that incorporating point cloud assistance into the compression pipeline consistently enhances the performance.

5.6CVJan 26, 2021Code
AINet+: Advancing Superpixel Segmentation via Cascaded Association Implantation

Yaxiong Wang, Yunchao Wei, Yujiao Wu et al.

Superpixel segmentation has seen significant progress benefiting from the deep convolutional networks. The typical approach entails initial division of the image into grids, followed by a learning process that assigns each pixel to adjacent grid segments. However, reliance on convolutions with confined receptive fields results in an implicit, rather than explicit, understanding of pixel-grid interactions. This limitation often leads to a deficit of contextual information during the mapping of associations. To counteract this, we introduce the Association Implantation (AI) module, designed to allow networks to explicitly engage with pixel-grid relationships. This module embeds grid features directly into the vicinity of the central pixel and employs convolutional operations on an enlarged window, facilitating an adaptive transfer of knowledge. This approach enables the network to explicitly extract context at the pixel-grid level, which is more aligned with the objectives of superpixel segmentation than mere pixel-wise interactions. By integrating the AI module across various layers, we enable a progressive refinement of pixel-superpixel relationships from coarse to fine. To further enhance the assignment of boundary pixels, we've engineered a boundary-aware loss function. This function aids in the discrimination of boundary-adjacent pixels at the feature level, thereby empowering subsequent modules to precisely identify boundary pixels and enhance overall boundary accuracy. Our method has been rigorously tested on four benchmarks, including BSDS500, NYUv2, ACDC, and ISIC2017, and our model can achieve competitive performance with comparison methods.

9.0CVNov 6, 2019
Localization-aware Channel Pruning for Object Detection

Zihao Xie, Wenbing Tao, Li Zhu et al.

Channel pruning is one of the important methods for deep model compression. Most of existing pruning methods mainly focus on classification. Few of them conduct systematic research on object detection. However, object detection is different from classification, which requires not only semantic information but also localization information. In this paper, based on discrimination-aware channel pruning (DCP) which is state-of-the-art pruning method for classification, we propose a localization-aware auxiliary network to find out the channels with key information for classification and regression so that we can conduct channel pruning directly for object detection, which saves lots of time and computing resources. In order to capture the localization information, we first design the auxiliary network with a contextual ROIAlign layer which can obtain precise localization information of the default boxes by pixel alignment and enlarges the receptive fields of the default boxes when pruning shallow layers. Then, we construct a loss function for object detection task which tends to keep the channels that contain the key information for classification and regression. Extensive experiments demonstrate the effectiveness of our method. On MS COCO, we prune 70\% parameters of the SSD based on ResNet-50 with modest accuracy drop, which outperforms the-state-of-art method.