SDAIJul 9

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction

arXiv:2607.081115.2h-index: 5
Predicted impact top 61% in SD · last 90 daysOriginality Incremental advance
AI Analysis

For researchers working on speaker extraction in real-world conversations, PS4 provides a practical training method without requiring clean target speech, achieving competitive results on a public benchmark.

PS4 introduces a proxy-supervised training framework for target speaker extraction in real conversational mixtures, using a large-scale corpus of 71,771 samples and four differentiable objectives. It achieves 2nd place on the REAL-T challenge leaderboard with the best speaker similarity and timing F1.

Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora and clean target speech for supervision are unavailable. We present PS4, a proxy-supervised training framework for TSE in real conversational mixtures, with two main contributions. First, we construct a large-scale corpus of 71,771 training samples derived from four public datasets, covering both Chinese and English scenarios. Each sample contains an overlapping speech mixture, per-speaker enrollment audio, a ground-truth transcript, and frame-level voice activity labels. Second, we propose a proxy-supervised joint training strategy that fine-tunes a BSRNN-based TSE model using four complementary differentiable objectives: ASR cross-entropy, speaker similarity, frame-level voice activity detection, and perceptual audio quality. Starting from a publicly available pre-trained checkpoint, only the BSRNN separator is updated during fine-tuning. On the REAL-T challenge leaderboard, PS4 ranks 2nd overall, achieving the best speaker similarity and timing F1 among all submitted systems.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes