A compact representation of visual speech data using latent variables.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 24231875.
- Also identified by DOI 10.1109/TPAMI.2013.173.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
The problem of visual speech recognition involves the decoding of the video dynamics of a talking mouth in a high-dimensional visual space. In this paper, we propose a generative latent variable model to provide a compact representation of visual speech data. The model uses latent variables to separately represent the interspeaker variations of visual appearances and those caused by uttering within images, and incorporates the structural information of the visual data through placing priors of the latent variables along a curve embedded within a path graph.
Medical subject headings
- Pattern Recognition, Automated
- Speech
- Speech Recognition Software
- Video Recording