Lightweight real-time speech enhancement: State-space models and multi-spectral scanning techniques.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40652906.
- Also identified by DOI 10.1016/j.neunet.2025.107805.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Achieving efficient, real-time speech enhancement requires a careful balance between signal quality and computational complexity. In this paper, we propose a lightweight end-to-end framework that leverages Mamba State-Space Model (SSM) in combination with multi-spectral scanning techniques to improve speech signals in real-time communication systems. To address the distinct challenges of speech processing, we introduce three novel spectrogram scanning methods designed to capture long-term dependencies across full-band, sub-band, and cross-band spectrums. These techniques are enhanced by ERB compression, which simulates human auditory perception, and high-frequency reconstruction to reduce computational overhead. Additionally, we incorporate a Voice Activity Detection (VAD) loss function to refine the recovery of coarse-grained temporal features during training. Our model is trained using a large, synthetic dataset generated through a custom pipeline and evaluated on the ICASSP SSI Challenge dataset. The experimental results demonstrate that our approach outperforms state-of-the-art models in terms of overall quality (OVRL) and speech signal clarity (SIG), while remaining highly efficient in terms of resource consumption. Besides, on the VoiceBank+DEMAND benchmark, our framework surpasses competitive models in overall speech enhancement, demonstrating the potential of our lightweight solution for practical applications.
Medical subject headings
- Speech
- Neural Networks, Computer
- Signal Processing, Computer-Assisted