SDJul 7

Flow Matching-Based Speech Source Separation with Best-of-N Biometric Sampling

arXiv:2607.060889.6
Predicted impact top 32% in SD · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and engineers deploying single-channel speech separation in real-world applications, this work offers a practical solution to permutation ambiguity and generative sampling variability, with demonstrated improvements in downstream tasks.

The paper introduces a flow-matching-based speech separation method that uses a frozen speaker encoder to resolve permutation ambiguity and enable best-of-N sampling for improved quality. The method achieves competitive SI-SDR, PESQ, and ESTOI scores on Libri2Mix and sets new state-of-the-art in downstream ASR (cpWER) and speaker verification (EER).

Single-channel speech separation remains challenging for real-world deployment due to source permutation ambiguity, sampling variability of generative models, and the difficulty of processing long recordings with chunk-wise inference. We address these issues with a conditional flow-matching-based method that produces an ordered two-source output conditioned on the mixture. A frozen speaker encoder defines the source order during training and is reused at inference for biometric best-of-$N$ candidate selection and chunk-level channel alignment. We evaluate separation quality on Libri2Mix benchmark using SI-SDR, PESQ, and ESTOI, and measure downstream impact using cpWER for automatic speech recognition and EER for speaker verification. The results show that the proposed Transformer U-Net variant is competitive with strong baselines in objective separation metrics and achieves the lowest downstream automatic speech recognition and speaker verification error rates in all evaluated settings.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes