Detecting change in clinical communication skills: Responsiveness of the Gap-Kalamazoo communication skills assessment form scored by experts and ChatGPT.

Wang, Yi-Ching; Lee, Ya-Chen; Ju, Yu-Jeng; Chen, Tzu-Ting; Huang, Sheau-Ling; Yang, Chih-Wei; Hsieh, Ching-Lin · Patient Educ Couns · 2026

other

Where this comes from

Abstract

The Gap-Kalamazoo Communication Skills Assessment Form (GKCSAF) is widely used to assess clinicians' communication skills. While expert-rated administration of the GKCSAF is seen as objective and thorough, manual scoring is time- and resource-intensive. The ChatGPT-rated GKCSAF is a large language model-based rating approach that uses the same GKCSAF items and scoring criteria. It shows acceptable reliability and validity while reducing rater burden. However, the responsiveness of the expert-rated and ChatGPT-rated GKCSAF has remained unknown, limiting the interpretation of change scores in communication skills. The study aimed to compare the responsiveness of the expert-rated and ChatGPT-rated GKCSAF. Eighty occupational therapy students completed two recorded clinical interactions with different patients. Between interactions, students received training via the Communication Skills Measure for Therapists, a formative measure that assesses communication skills and yields actionable feedback. Transcripts of the interactions were independently scored by trained expert raters and by ChatGPT. Responsiveness was examined using paired t-tests, Cohen's d, and standardized response mean (SRM). Differences in responsiveness indices between the rating approaches (i.e., the expert-rated and ChatGPT-rated GKCSAF) were estimated using bootstrap confidence intervals. For both rating approaches, mean post-training scores were significantly higher than mean pre-training scores (both t = 3.8, p < 0.001). The responsiveness indices were nearly moderate (expert-rated GKCSAF: d = 0.45, SRM = 0.42; ChatGPT-rated GKCSAF: d = 0.47, SRM = 0.43), with no significant differences between the two rating approaches (bootstrap 95% confidence intervals for the differences included zero). Both the ChatGPT-rated and expert-rated GKCSAF showed evidence of responsiveness for detecting changes in clinicians' communication skills. Considering the shorter scoring time and reduced rater burden of the ChatGPT-rated GKCSAF, it may be a promising outcome assessment for clinical and research use.