Neuromorphic Speech Enhancement with Dual-Branch Spiking Neural Networks
This work advances neuromorphic speech enhancement by demonstrating competitive performance with far fewer parameters, addressing the efficiency-accuracy gap for energy-constrained applications.
The authors propose a dual-branch spiking neural network (GSU-DBNet) for speech enhancement that jointly models magnitude and complex spectra, achieving a PESQ of 3.04 with only 394K parameters, outperforming prior SNN methods and using 4.5%–10.6% of the parameters of ANN-based models.
Spiking neural network (SNN)-based neuromorphic speech enhancement has emerged as a promising paradigm due to its energy efficiency, yet it still underperforms classical artificial neural network (ANN)-based approaches owing to binary activations and the lack of well-designed network architectures. To overcome this limitation, we propose a novel dual-branch spiking neural network architecture equipped with a gated spiking unit (GSU), termed GSU-DBNet. Specifically, GSU-DBNet simultaneously models the speech magnitude spectrum and complex spectrum, predicting the corresponding magnitude and complex spectral masks. Meanwhile, a dual-path GSU module is adopted to exploit temporal and frequency information for enhanced spatiotemporal feature representation. Experiments on a popular benchmark dataset show that GSU-DBNet achieves a PESQ score of 3.04 with only 394K parameters, outperforming existing SNN-based methods while using only 4.5%--10.6% of the parameters of representative ANN-based models.