Lightweight real-time speech enhancement: State-space models and multi-spectral scanning techniques.

Zhu, Xiaodong; Yang, Junqi; Yang, Yuhong; Tu, Weiping; Wang, Zhongyuan · Neural Netw · 2025

basic_science · Level V

Where this comes from

Abstract

Achieving efficient, real-time speech enhancement requires a careful balance between signal quality and computational complexity. In this paper, we propose a lightweight end-to-end framework that leverages Mamba State-Space Model (SSM) in combination with multi-spectral scanning techniques to improve speech signals in real-time communication systems. To address the distinct challenges of speech processing, we introduce three novel spectrogram scanning methods designed to capture long-term dependencies across full-band, sub-band, and cross-band spectrums. These techniques are enhanced by ERB compression, which simulates human auditory perception, and high-frequency reconstruction to reduce computational overhead. Additionally, we incorporate a Voice Activity Detection (VAD) loss function to refine the recovery of coarse-grained temporal features during training. Our model is trained using a large, synthetic dataset generated through a custom pipeline and evaluated on the ICASSP SSI Challenge dataset. The experimental results demonstrate that our approach outperforms state-of-the-art models in terms of overall quality (OVRL) and speech signal clarity (SIG), while remaining highly efficient in terms of resource consumption. Besides, on the VoiceBank+DEMAND benchmark, our framework surpasses competitive models in overall speech enhancement, demonstrating the potential of our lightweight solution for practical applications.

Medical subject headings