Spectro-temporal vs. spectral features to predict the lombard gain in Mandarin Chinese.
other
Where this comes from
- Record sourced from PubMed, PMID 42647518.
- Also identified by DOI 10.1371/journal.pone.0356236.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
In background noise, speakers adapt their speech production, giving rise to Lombard speech, which often improves speech intelligibility (SI). While intelligibility benefits of Lombard speech have been extensively studied in non-tonal languages, it remains unclear whether spectro-temporal cues, which are critical for tonal contrasts, are necessary to predict both the Lombard gain (LG) (i.e., intelligibility improvement relative to plain speech) and absolute SI in Mandarin Chinese. Predictions of two SI-models were compared, namely, an automatic speech recognition (ASR)-based approach using spectral or spectro-temporal features and the speech intelligibility index (SII)-based model using spectral features. Predicted LG and absolute speech recognition threshold (SRT) values, for five female and six male speakers in stationary speech-shaped noise, were compared with empirical data. For both models, spectral features alone are sufficient for accurate prediction of the LG for both models. In contrast, predictions of absolute SRT were most accurate when spectro-temporal features were included, capturing substantial inter-speaker variability. Despite the tonal nature of Mandarin, spectro-temporal features are not required to predict the LG. However, they are essential to predict the absolute SRTs, which vary across speakers.
Medical subject headings
- Language
- Speech Intelligibility
- Speech Perception