Artificial intelligence-based automated scoring of the Clock Drawing Test in older adults in primary care: association with Mini-Mental State Examination scores, discriminative performance, and agreement with clinician rating.

Fam Pract · 2026

cross_sectional · Level IV

Where this comes from

Abstract

BACKGROUND: Cognitive impairment is common in older adults, and its early detection in primary care remains challenging. The Clock Drawing Test (CDT) is a practical tool, but variability in its scoring limits standardization. OBJECTIVE: To evaluate artificial intelligence (AI)-based automated CDT scoring in primary care by examining its association with Mini-Mental State Examination (MMSE) scores, discriminative performance for cognitive impairment, and agreement with clinician rating. METHODS: In this cross-sectional study, 207 adults aged ≥65 years were assessed in a primary care setting. CDT drawings were scored manually by a neurologist and automatically using a multimodal generative AI system based on the Manos and Wu 10-point method. Associations with MMSE were analyzed, and discriminative performance was assessed using receiver operating characteristic analysis. Agreement was evaluated using intraclass correlation coefficient and Bland-Altman analysis. RESULTS: AI-based CDT scores demonstrated a moderate positive correlation with MMSE (ρ = .437, P < .001). Both methods significantly discriminated MMSE-defined cognitive impairment, although manual scoring yielded a significantly higher area under the receiver operating characteristic curve than AI-based scoring (.779 vs .715; DeLong P = .028). AI-based scoring provided higher sensitivity but lower specificity. Agreement between methods was good (intraclass correlation coefficient = .721), with AI tending to assign slightly higher scores (mean difference: 1.13). CONCLUSION: AI-based CDT scoring is associated with global cognitive performance and offers meaningful discrimination of cognitive impairment in primary care. However, given lower accuracy and individual-level variability, it should be considered a supportive tool rather than a replacement for clinician assessment.