Haofeng Chen

CV
h-index6
4papers
3,158citations
Novelty40%
AI Score31

4 Papers

24.8CVOct 12, 2022
QDTrack: Quasi-Dense Similarity Learning for Appearance-Only Multiple Object Tracking

Tobias Fischer, Thomas E. Huang, Jiangmiao Pang et al. · eth-zurich, mit

Similarity learning has been recognized as a crucial step for object tracking. However, existing multiple object tracking methods only use sparse ground truth matching as the training objective, while ignoring the majority of the informative regions in images. In this paper, we present Quasi-Dense Similarity Learning, which densely samples hundreds of object regions on a pair of images for contrastive learning. We combine this similarity learning with multiple existing object detectors to build Quasi-Dense Tracking (QDTrack), which does not require displacement regression or motion priors. We find that the resulting distinctive feature space admits a simple nearest neighbor search at inference time for object association. In addition, we show that our similarity learning scheme is not limited to video data, but can learn effective instance similarity even from static input, enabling a competitive tracking performance without training on videos or using tracking supervision. We conduct extensive experiments on a wide variety of popular MOT benchmarks. We find that, despite its simplicity, QDTrack rivals the performance of state-of-the-art tracking methods on all benchmarks and sets a new state-of-the-art on the large-scale BDD100K MOT benchmark, while introducing negligible computational overhead to the detector.

22.4CVMay 11, 2021Code
Home Action Genome: Cooperative Compositional Action Understanding

Nishant Rai, Haofeng Chen, Jingwei Ji et al.

Existing research on action recognition treats activities as monolithic events occurring in videos. Recently, the benefits of formulating actions as a combination of atomic-actions have shown promise in improving action understanding with the emergence of datasets containing such annotations, allowing us to learn representations capturing this information. However, there remains a lack of studies that extend action composition and leverage multiple viewpoints and multiple modalities of data for representation learning. To promote research in this direction, we introduce Home Action Genome (HOMAGE): a multi-view action dataset with multiple modalities and view-points supplemented with hierarchical activity and atomic action labels together with dense scene composition labels. Leveraging rich multi-modal and multi-view settings, we propose Cooperative Compositional Action Understanding (CCAU), a cooperative learning framework for hierarchical action recognition that is aware of compositional action elements. CCAU shows consistent performance improvements across all modalities. Furthermore, we demonstrate the utility of co-learning compositions in few-shot action recognition by achieving 28.6% mAP with just a single sample.

52.5CVMay 12, 2018Code
BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning

Fisher Yu, Haofeng Chen, Xin Wang et al.

Datasets drive vision progress, yet existing driving datasets are impoverished in terms of visual content and supported tasks to study multitask learning for autonomous driving. Researchers are usually constrained to study a small set of problems on one dataset, while real-world computer vision applications require performing tasks of various complexities. We construct BDD100K, the largest driving video dataset with 100K videos and 10 tasks to evaluate the exciting progress of image recognition algorithms on autonomous driving. The dataset possesses geographic, environmental, and weather diversity, which is useful for training models that are less likely to be surprised by new conditions. Based on this diverse dataset, we build a benchmark for heterogeneous multitask learning and study how to solve the tasks together. Our experiments show that special training strategies are needed for existing models to perform such heterogeneous tasks. BDD100K opens the door for future studies in this important venue.

3.2RONov 4, 2017
Analyzing Material Recognition Performance of Thermal Tactile Sensing using a Large Materials Database and a Real Robot

Haoping Bai, Haofeng Chen, Elizabeth Healy et al.

In this paper we focus on analyzing the thermal modality of tactile sensing for material recognition using a large materials database. Many factors affect thermal recognition performance, including sensor noise, the initial temperatures of the sensor and the object, the thermal effusivities of the materials, and the duration of contact. To analyze the influence of these factors on thermal recognition, we used a semi-infinite solid based thermal model to simulate heat-transfer data from all the materials in the CES Edupack Level-1 database. We used support-vector machines (SVMs) to predict F1 scores for binary material recognition for 2346 material pairs. We also collected data using a real robot equipped with a thermal sensor and analyzed its material recognition performance on 66 real-world material pairs. Additionally, we analyzed the performance when the models were trained on the simulated data and tested on the real-robot data. Our models predicted the material recognition performance with a 0.980 F1 score for the simulated data, a 0.994 F1 score for real-world data with constant initial sensor temperatures, a 0.966 F1 score for real-world data with varied initial sensor temperatures, and a 0.815 F1 score for sim-to-real transfer. Finally, we present some guidelines on sensor design and parameter choice for thermal recognition based on the insights gained from these results that would hopefully enable robotics researchers to use this less-explored tactile sensing modality more effectively during physical human-robot and robot-object interactions. We release our simulated and real-robot datasets for further use by the robotics community.