3.3ASMay 24, 2025
TS-URGENet: A Three-stage Universal Robust and Generalizable Speech Enhancement NetworkXiaobin Rong, Dahan Wang, Qinwen Hu et al.
Universal speech enhancement aims to handle input speech with different distortions and input formats. To tackle this challenge, we present TS-URGENet, a Three-Stage Universal, Robust, and Generalizable speech Enhancement Network. To address various distortions, the proposed system employs a novel three-stage architecture consisting of a filling stage, a separation stage, and a restoration stage. The filling stage mitigates packet loss by preliminarily filling lost regions under noise interference, ensuring signal continuity. The separation stage suppresses noise, reverberation, and clipping distortion to improve speech clarity. Finally, the restoration stage compensates for bandwidth limitation, codec artifacts, and residual packet loss distortion, refining the overall speech quality. Our proposed TS-URGENet achieved outstanding performance in the Interspeech 2025 URGENT Challenge, ranking 2nd in Track 1.
DPCRN: Dual-Path Convolution Recurrent Network for Single Channel Speech EnhancementXiaohuai Le, Hongsheng Chen, Kai Chen et al.
The dual-path RNN (DPRNN) was proposed to more effectively model extremely long sequences for speech separation in the time domain. By splitting long sequences to smaller chunks and applying intra-chunk and inter-chunk RNNs, the DPRNN reached promising performance in speech separation with a limited model size. In this paper, we combine the DPRNN module with Convolution Recurrent Network (CRN) and design a model called Dual-Path Convolution Recurrent Network (DPCRN) for speech enhancement in the time-frequency domain. We replace the RNNs in the CRN with DPRNN modules, where the intra-chunk RNNs are used to model the spectrum pattern in a single frame and the inter-chunk RNNs are used to model the dependence between consecutive frames. With only 0.8M parameters, the submitted DPCRN model achieves an overall mean opinion score (MOS) of 3.57 in the wide band scenario track of the Interspeech 2021 Deep Noise Suppression (DNS) challenge. Evaluations on some other test sets also show the efficacy of our model.
1.9SDAug 1, 2020
Efficient Independent Vector Extraction of Dominant Target SpeechLele Liao, Zhaoyi Gu, Jing Lu
The complete decomposition performed by blind source separation is computationally demanding and superfluous when only the speech of one specific target speaker is desired. In this paper, we propose a computationally efficient blind speech extraction method based on a proper modification of the commonly utilized independent vector analysis algorithm, under the mild assumption that the average power of signal of interest outweighs interfering speech sources. Considering that the minimum distortion principle cannot be implemented since the full demixing matrix is not available, we also design a one-unit scaling operation to solve the scaling ambiguity. Simulations validate the efficacy of the proposed method in extracting the dominant speech.
Nonlinear Residual Echo Suppression Based on Multi-stream Conv-TasNetHongsheng Chen, Teng Xiang, Kai Chen et al.
Acoustic echo cannot be entirely removed by linear adaptive filters due to the nonlinear relationship between the echo and far-end signal. Usually a post processing module is required to further suppress the echo. In this paper, we propose a residual echo suppression method based on the modification of fully convolutional time-domain audio separation network (Conv-TasNet). Both the residual signal of the linear acoustic echo cancellation system, and the output of the adaptive filter are adopted to form multiple streams for the Conv-TasNet, resulting in more effective echo suppression while keeping a lower latency of the whole system. Simulation results validate the efficacy of the proposed method in both single-talk and double-talk situations.
1.2ASApr 14, 2019
A robust DOA estimation method for a linear microphone array under reverberant and noisy environmentsHao Wang, Jing Lu
A robust method for linear array is proposed to address the difficulty of direction-of-arrival (DOA) estimation in reverberant and noisy environments. A direct-path dominance test based on the onset detection is utilized to extract time-frequency bins containing the direct propagation of the speech. The influence of the transient noise, which severely contaminates the onset test, is mitigated by a proper transient noise determination scheme. Then for voice features, a two-stage procedure is designed based on the extracted bins and an effective dereverberation method, with robust but possibly biased estimation from middle frequency bins followed by further refinement in higher frequency bins. The proposed method effectively alleviates the estimation bias caused by the linear arrangement of microphones, and has stable performance under noisy and reverberant environments. Experimental evaluation using a 4-element microphone array demonstrates the efficacy of the proposed method.
1.2ASFeb 25, 2018
Frequency domain TRINICON-based blind source separation method with multi-source activity detection for sparsely mixed signalsZelin Wang, Jing Lu, Kai chen
The TRINICON ('Triple-N ICA for convolutive mixtures') framework is an effective blind signal separation (BSS) method for separating sound sources from convolutive mixtures. It makes full use of the non-whiteness, non-stationarity and non-Gaussianity properties of the source signals and can be implemented either in time domain or in frequency domain, avoiding the notorious internal permutation problem. It usually has best performance when the sources are continuously mixed. In this paper, the offline dual-channel frequency domain TRINICON implementation for sparsely mixed signals is investigated, and a multi-source activity detection is proposed to locate the active period of each source, based on which the filter updating strategy is regularized to improve the separation performance. The objective metric provided by the BSSEVAL toolkit is utilized to evaluate the performance of the proposed scheme.
3.3ASFeb 25, 2018
RLS-Based Adaptive Dereverberation Tracing Abrupt Position Change of Target SpeakerTeng Xiang, Jing Lu, Kai Chen
Adaptive algorithm based on multi-channel linear prediction is an effective dereverberation method balancing well between the attenuation of the long-term reverberation and the dereverberated speech quality. However, the abrupt change of the speech source position, usually caused by the shift of the speakers, forms an obstacle to the adaptive algorithm and makes it difficult to guarantee both the fast convergence speed and the optimal steady-state behavior. In this paper, the RLS-based adaptive multi-channel linear prediction method is investigated and a time-varying forgetting factor based on the relative weighted change of the adaptive filter coefficients is proposed to effectively tracing the abrupt change of the target speaker position. The advantages of the proposed scheme are demonstrated in the simulations and experiments.