SDAIJun 28

TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation

arXiv:2606.295754.6
Predicted impact top 66% in SD · last 90 daysOriginality Incremental advance
AI Analysis

For researchers deploying speech separation on edge devices, TF-MoE offers a way to improve performance without increasing computational cost.

TF-MoE introduces a sparse Mixture-of-Experts framework for speech separation that improves model capacity without increasing inference cost, achieving +3.8 dB SDR over BSRNN on Libri2Mix at comparable computational cost.

Recent advances in speech separation (SS) have led to compact front-end models with small parameter sizes, yet their high computational cost remains a major barrier for deployment on edge devices. To address this, we propose TF-MoE, a sparse Mixture-of-Experts (MoE) framework that enhances model capacity with almost no increase in inference cost. Our method introduces dynamic expert specialization in time and frequency dimensions through alternating time-wise and frequency-wise MoE modules, each dynamically selecting experts per frame or mel band. Built upon a mel-band-splitting Conformer backbone, TF-MoE achieves strong performance on SS tasks under low-compute settings. Experimental results demonstrate that TF-MoE consistently improves separation performance under computation cost constraints, outperforming BSRNN by +3.8 dB SDR on Libri2Mix with comparable 4.1 GMACs/s inference cost. This positions TF-MoE as a promising candidate for edge-device deployment.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes