ChatGPT versus expert arthroplasty surgeons in total knee arthroplasty patient counseling.

Liu, Jonathan; Daher, Mohammad; Laperche, Jacob; Gilreath, Noah; Testa, Edward J; El-Othamni, Mouhanad M; Barrett, Thomas J; Antoci, Valentin · Knee · 2025

cross_sectional · Level IV

Where this comes from

Abstract

This study aimed to assess the effectiveness of AI compared directly with expert arthroplasty surgeons regarding patient counseling for total knee arthroplasty (TKA). A set of 10 commonly asked generic and nonspecific, single-step patient questions were selected based on review of existing patient resources and expert consensus. Responses were then collected from ChatGPT-4.0 as well as five expert arthroplasty attendings at our institution. A, B, C, D, and E represent attending responses, while F represents the ChatGPT responses. The collected responses were then blinded and independently assessed by the same five arthroplasty surgeons using a five-point Likert scale in four performance areas including empathy, accuracy, completeness, and overall quality. Average scores for each question were determined. Set F, the ChatGPT answers scored significantly higher than sets A, B, and D in all categories. However, set F did not differ significantly from set C, and E in all the categories. The mean score for set D was above a mean of 4, above neutral, for all four categories. This was only the case for sets C and E.When the attendings scores were combined and compared with ChatGPT, the latter had higher ratings for empathy (4.4 vs. 3.5), accuracy (4.4 vs. 3.7), completeness (4.4 vs. 3.5), and overall quality (4.4 vs. 3.6) (P < 0.001). A preliminary evaluation of ChatGPT-4.0 shows potential for large language AI models to serve as a supplementary resource of patients considering TKA.

Medical subject headings

Anatomy