Enhancing end-to-end speech translation via multi-stage knowledge distillation.
Where this comes from
- Record sourced from PubMed, PMID 41429075.
- Also identified by DOI 10.1016/j.neunet.2025.108444.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Knowledge distillation (KD) using machine translation (MT) teacher models has demonstrated effectiveness in enhancing end-to-end speech-to-text translation (ST). However, existing KD methods for ST primarily rely on the teacher's output distributions, transferring only partial knowledge and failing to capture the teacher model's deep representation capabilities. Consequently, the student model struggles to fully acquire translation knowledge, degrading translation quality. Moreover, these methods depend on limited ST data during the distillation process. The scarcity of ST data hinders the comprehensive translation knowledge transfer from the MT model to the ST model, further limiting their potential. In this paper, we propose a multi-grained distillation method to enhance knowledge transfer. Specifically, in addition to conventional sequence-level distillation, we introduce adaptive word-level distillation, which dynamically adjusts word-level weights to prioritize challenging translations, and cross-modal hidden state distillation, which aligns the hidden states of the MT and ST models to bridge the modality gap between speech and text, facilitating more effective cross-modal translation knowledge transfer. Building on this, we design a multi-stage knowledge distillation (MSKD) framework that systematically leverages large-scale external automatic speech recognition (ASR) and MT data, ST task data, and multi-grained distillation to progressively improve ST performance, ensuring comprehensive utilization of external resources and effective translation knowledge transfer. Extensive experiments demonstrate that MSKD achieves state-of-the-art performance in scenarios using only the MuST-C dataset, significantly outperforming previous distillation methods for ST (+2.5 BLEU). With external data, MSKD surpasses strong end-to-end ST baselines (+2.9 BLEU) and cascaded systems (+1.9 BLEU), highlighting its effectiveness and scalability.
Medical subject headings
- Distillation
- Speech
- Knowledge
- Translating
- Speech Recognition Software
- Machine Learning