Evaluating the Evolution of ChatGPT as an Information Resource in Shoulder and Elbow Surgery.

Nieves-Lopez, Benjamin; Bechtle, Alexandra R; Traverse, Jennifer; Klifto, Christopher; Schoch, Bradley S; Aziz, Keith T · Orthopedics · 2025

cross_sectional · Level IV

Where this comes from

Abstract

The purpose of this study was to evaluate the performance and evolution of Chat Generative Pre-Trained Transformer (ChatGPT; OpenAI) as a resource for shoulder and elbow surgery information by assessing its accuracy on the American Academy of Orthopaedic Surgeons shoulder-elbow self-assessment questions. We hypothesized that both ChatGPT models would demonstrate proficiency and that there would be significant improvement with progressive iterations. A total of 200 questions were selected from the 2019 and 2021 American Academy of Orthopaedic Surgeons shoulder-elbow self-assessment questions. ChatGPT 3.5 and 4 were used to evaluate all questions. Questions with non-text data were excluded (114 questions). Remaining questions were input into ChatGPT and categorized as follows: anatomy, arthroplasty, basic science, instability, miscellaneous, nonoperative, and trauma. ChatGPT's performances were quantified and compared across categories with chi-square tests. The continuing medical education credit threshold of 50% was used to determine proficiency. Statistical significance was set at <i>P</i><.05. ChatGPT 3.5 and 4 answered 52.3% and 73.3% of the questions correctly, respectively (<i>P</i>=.003). ChatGPT 3.5 performed significantly better in the instability category (<i>P</i>=.037). ChatGPT 4's performance did not significantly differ across categories (<i>P</i>=.841). ChatGPT 4 performed significantly better than ChatGPT 3.5 in all categories except instability and miscellaneous. ChatGPT 3.5 and 4 exceeded the proficiency threshold. ChatGPT 4 performed better than ChatGPT 3.5, showing an increased capability to correctly answer shoulder and elbow-focused questions. Further refinement of ChatGPT's training may improve its performance and utility as a resource. Currently, ChatGPT remains unable to answer questions at a high enough accuracy to replace clinical decision-making. [<i>Orthopedics</i>. 2025;48(2):e69-e74.].

Medical subject headings

Anatomy