ASSDJun 23

Joint Learning of Covariance Estimation and White Noise Gain for Robust MVDR Beamforming

arXiv:2606.241374.7
Predicted impact top 84% in AS · last 90 daysOriginality Incremental advance
AI Analysis

For practitioners of multichannel speech enhancement, this work addresses the sensitivity of MVDR beamformers to microphone self-noise and array mismatches by adaptively estimating the WNG constraint, improving performance under unknown or time-varying acoustic conditions.

This paper proposes a data-driven MVDR beamforming framework that uses a deep neural network to jointly estimate a time-frequency noise mask for covariance estimation and a frequency-dependent white noise gain (WNG) threshold, enabling adaptive robustness-directivity control. Experiments show consistent improvements in speech quality and intelligibility over conventional fixed-WNG MVDR methods.

The minimum variance distortionless response (MVDR) beamformer is widely used for multichannel speech enhancement due to strong noise suppression while preserving target signals. In practice, its performance is sensitive to microphone self-noise and array mismatches. Existing approaches typically rely on fixed, manually tuned WNG thresholds or diagonal loading, leading to suboptimal performance under unknown or time-varying acoustic conditions. This paper proposes a data-driven MVDR framework that adaptively estimates the WNG constraint using a deep neural network. The network jointly predicts a time-frequency noise mask for covariance estimation and a frequency-dependent WNG threshold, enabling dynamic robustness-directivity control. A differentiable robust MVDR layer is integrated into the framework, allowing end-to-end optimization. Experiments demonstrate consistent improvements in speech quality and intelligibility over conventional fixed-WNG MVDR methods.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes