SDJul 9

It Takes Few to TANGO: A Quantized Distributed Model for Binaural Speech Enhancement

arXiv:2607.086452.1
Predicted impact top 90% in SD · last 90 daysOriginality Synthesis-oriented
AI Analysis

For deploying speech enhancement on resource-constrained devices, this work provides a quantized model that reduces computational and memory requirements without significant performance loss.

This paper investigates low-precision inference for TANGO, a hybrid distributed binaural speech enhancement system, and shows that quantization errors in mask estimates are compensated by spatial filtering, enabling a compact model (MN-TANGO) with 4.65 MMAC/s and 0.177 MB while maintaining performance.

Neural network-based multichannel speech enhancement systems achieve strong enhancement performance, but their computational and memory requirements limit deployment on resource-constrained devices. This paper investigates low-precision inference for TANGO, a hybrid distributed binaural speech enhancement system combining neural mask estimation with spatial filtering. We evaluate post-training quantization and quantization-aware training for the neural components, and analyze how quantization errors in the mask estimators propagate through the downstream spatial filtering stage. Our analysis shows that, although quantization degrades intermediate mask estimates, the spatial filtering stage compensates for most quantization-induced errors. Leveraging this robustness, we simplify TANGO into MN-TANGO, reducing both model size and computational complexity while maintaining comparable final performance. By combining INT8 weight-and-activation quantization with ERB compression and grouped recurrent layers, the most compact MN-TANGO reaches 4.65 MMAC/s and 0.177 MB.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes