SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
This challenge provides a realistic benchmark for target speaker extraction, addressing the gap between simulated and real-world conditions for researchers working on speech separation and extraction.
The REAL-TSE Challenge introduces a benchmark for target speaker extraction from real conversational recordings with natural overlap, reverberation, and noise, evaluating systems on Mandarin and English data. The challenge includes online and offline tracks, with submitted systems showing that offline processing outperforms online, and that enrollment quality significantly impacts performance.
We introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and one or more enrollment utterances from a target speaker, participating systems must recover only the target speech. Unlike simulated read-speech benchmarks, REAL-TSE evaluates Mandarin and English recordings that contain natural overlap, reverberation, noise, channel mismatch, and conversational dynamics. The challenge defines two complementary tracks: an Online track for low-latency streaming extraction and an Offline track for full-context processing. Systems are evaluated with Token Error Rate (TER), Speaker Similarity (SpkSim), DNSMOS, and target-speaker activity F1. This overview paper describes the task definition, datasets, baselines, evaluation protocol, submitted systems, condition-wise findings, and lessons for future real-world TSE benchmarks.