MsDUNE: A multi-scale masked temporal fusion framework for speaker-independent lipreading via Dirichlet uncertainty estimation.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40618471.
- Also identified by DOI 10.1016/j.neunet.2025.107783.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Lipreading, the task of recognizing speech based on visual cues from lip movements, typically requires a substantial amount of labeled training data to achieve optimal performance. However, this task is highly sensitive to variations among speakers, often resulting in significantly degraded recognition accuracy for unseen speakers. In this work, we introduce a novel framework, multi-scale masked temporal fusion with Dirichlet uncertainty estimation (MsDUNE), designed to mitigate the feature distribution disparities across different speakers. The proposed framework leverages a Dirichlet distribution to parameterize the latent space of a single feature branch, which is then quantitatively assessed through evidence and belief masses. Furthermore, MsDUNE calibrates multi-scale feature distributions by accounting for the mutual influence of feature beliefs between two branches, thereby enhancing the generalization capability of the lipreading model. We validate our approach through extensive experiments conducted on two widely recognized benchmarks, LRW-ID and AV Letters, as well as a self-collected lipreading dataset, CVSR100. The experimental results highlight the state-of-the-art performance of our method, particularly in scenarios involving unseen or overlapping speakers.
Medical subject headings
- Lipreading