Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?
For developers of full-duplex dialogue systems, this work identifies a critical gap in controllable turn management, which is essential for real-world deployment across diverse applications.
The paper introduces Instruct-FD, a benchmark to evaluate whether full-duplex spoken dialogue systems can follow explicit turn-taking instructions. Benchmarking six state-of-the-art systems shows the best model achieves only 64.4% adherence, with proactive behaviors like interruption being especially challenging.
Current full-duplex (FD) spoken dialogue systems can produce fluid interactions, yet it remains unclear whether they can adapt their turn-taking behavior when explicitly instructed. This is critical for real-world deployment, where conversational policies vary across applications (e.g., proactive tutoring vs. passive counseling). We introduce Instruct-FD, an instruction-conditioned benchmark for evaluating controllable turn management in FD systems. To enable this, we develop a human-validated, scalable synthetic pipeline that generates instruction-conditioned conversations, along with a deployment-agnostic multi-turn evaluation protocol and an LLM-based judge. Benchmarking six state-of-the-art full-duplex systems reveals a substantial gap in instruction-following turn management: the best model achieves only 64.4% adherence. Performance is highly uneven across behaviors and scenarios, with proactive behaviors such as model backchanneling and interruption remaining particularly challenging. These findings establish instruction-following turn management as a crucial direction for building adaptable and deployable full-duplex dialogue systems.