6.6MMJun 25, 2021Code
Cross-Modal Knowledge Distillation Method for Automatic Cued Speech RecognitionJianrong Wang, Ziyue Tang, Xuewei Li et al.
Cued Speech (CS) is a visual communication system for the deaf or hearing impaired people. It combines lip movements with hand cues to obtain a complete phonetic repertoire. Current deep learning based methods on automatic CS recognition suffer from a common problem, which is the data scarcity. Until now, there are only two public single speaker datasets for French (238 sentences) and British English (97 sentences). In this work, we propose a cross-modal knowledge distillation method with teacher-student structure, which transfers audio speech information to CS to overcome the limited data problem. Firstly, we pretrain a teacher model for CS recognition with a large amount of open source audio speech data, and simultaneously pretrain the feature extractors for lips and hands using CS data. Then, we distill the knowledge from teacher model to the student model with frame-level and sequence-level distillation strategies. Importantly, for frame-level, we exploit multi-task learning to weigh losses automatically, to obtain the balance coefficient. Besides, we establish a five-speaker British English CS dataset for the first time. The proposed method is evaluated on French and British English CS datasets, showing superior CS recognition performance to the state-of-the-art (SOTA) by a large margin.
2.9ROFeb 24, 2018
Euler angles based loss function for camera relocalization with Deep learningQiang Fang, Tianjiang Hu
Deep learning has been applied to camera relocalization, in particular, PoseNet and its extended work are the convolutional neural networks which regress the camera pose from a single image. However there are many problems, one of them is expensive parameter selection. In this paper, we directly explore the three Euler angles as the orientation representation in the camera pose regressor. There is no need to select the parameter, which is not tolerant in the previous works. Experimental results on the 7 Scenes datasets and the King's College dataset demonstrate that it has competitive performances.
2.1RONov 10, 2016
Vision-aided Localization and Navigation Based on Trifocal TensorQiang Fang
In this paper, a novel method for vision-aided navigation based on trifocal tensor is presented. The main goal of the proposed method is to provide position estimation in GPS-denied environments for vehicles equipped with a standard inertial navigation systems(INS) and a single camera only. We treat the trifocal tensor as the measurement model, being only concerned about the vehicle state and do not estimate the the position of the tracked landmarks. The performance of the proposed method is demonstrated using simulation and experimental data.
2.2ROJul 26, 2012
A Unified Approach of Observability Analysis for Airborne SLAMQiang Fang, Xinsheng Huang
Observability is a key aspect of the state estimation problem of SLAM, However, the dimension and variables of SLAM system might be changed with new features, to which little attention is paid in the previous work. In this paper, a unified approach of observability analysis for SLAM system is provided, whether the dimension and variables of SLAM system are changed or not, we can use this approach to analyze the local or total observability of the SLAM system.