Na Li

h-index17
2papers
1,431citations

2 Papers

5.1IVAug 28, 2019
On Energy Compaction of 2D Saab Image Transforms

Na Li, Yongfei Zhang, Yun Zhang et al.

The block Discrete Cosine Transform (DCT) is commonly used in image and video compression due to its good energy compaction property. The Saab transform was recently proposed as an effective signal transform for image understanding. In this work, we study the energy compaction property of the Saab transform in the context of intra-coding of the High Efficiency Video Coding (HEVC) standard. We compare the energy compaction property of the Saab transform, the DCT, and the Karhunen-Loeve transform (KLT) by applying them to different sizes of intra-predicted residual blocks in HEVC. The basis functions of the Saab transform are visualized. Extensive experimental results are given to demonstrate the energy compaction capability of the Saab transform.

3.4LGMay 16, 2019
Learning discriminative features in sequence training without requiring framewise labelled data

Jun Wang, Dan Su, Jie Chen et al.

In this work, we try to answer two questions: Can deeply learned features with discriminative power benefit an ASR system's robustness to acoustic variability? And how to learn them without requiring framewise labelled sequence training data? As existing methods usually require knowing where the labels occur in the input sequence, they have so far been limited to many real-world sequence learning tasks. We propose a novel method which simultaneously models both the sequence discriminative training and the feature discriminative learning within a single network architecture, so that it can learn discriminative deep features in sequence training that obviates the need for presegmented training data. Our experiment in a realistic industrial ASR task shows that, without requiring any specific fine-tuning or additional complexity, our proposed models have consistently outperformed state-of-the-art models and significantly reduced Word Error Rate (WER) under all test conditions, and especially with highest improvements under unseen noise conditions, by relative 12.94%, 8.66% and 5.80%, showing our proposed models can generalize better to acoustic variability.