A 60-second interpretable voice model for early dementia screening.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 42507664.
- Also identified by DOI 10.1371/journal.pdig.0001552 and PMC identifier 13405093.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Early detection of cognitive impairment in assisted living is hindered by time-intensive tools like the Mini-Mental State Examination (MMSE) and the Montreal Cognitive Assessment (MoCA). We present a 60-second voice-based screening model that analyzes picture descriptions to estimate dementia risk. While recent deep learning approaches have shown promise on similar tasks, their lack of interpretability, large data requirements, and computational complexity limit clinical adoption. Using transcripts from the DementiaBank corpus, our model integrates traditional linguistic features (pause rate, pronoun use, syntactic complexity) with latent semantic dimensions extracted from language model embeddings. These semantic axes, interpretable constructs like "Drift & Hesitation" or "Over-detailed Narration", consistently emerged as top predictors and may represent novel linguistic biomarkers of early decline. The final ElasticNet classifier is sparse, interpretable, and outperforms known non-deep learning baselines (AUC = 0.858), exceeding MMSE. Its simplicity enables deployment in mobile apps or in-room monitors, offering scalable, low-burden screening for early dementia. This work supports a shift toward linguistically grounded, tech-enabled cognitive care in aging populations.