A scoping review: how evaluation methods shape our understanding of ChatGPT's effectiveness in healthcare.

Liu, Yuanyuan; Zhang, Yu; Mao, Haoran · Int J Med Inform · 2026

systematic_review · Level I

Where this comes from

Abstract

The rapid growth in research on ChatGPT's healthcare applications has led to diverse evaluation methods and substantially heterogeneous findings, undermining evidence reliability and hindering clinical translation. This review aims to examine how different evaluation methods shape our understanding of ChatGPT's effectiveness in healthcare. Studies published between 2023 and 2024 that assess the use of ChatGPT in medical or healthcare-related contexts were included. Evidence was obtained from peer-reviewed literature analyzing ChatGPT's applications across clinical, educational, and diagnostic domains. Following the PRISMA guidelines, this systematic review analyzed 131 studies published during 2023-2024 that assess the use of ChatGPT in medical contexts. The results indicate that predominant evaluation approaches-controlled trial studies, expert assessment studies, measurement-based evaluation studies, and prompt generation analysis studies-systematically influence conclusions about ChatGPT's performance due to their inherent methodological characteristics, such as subjectivity, objectivity, and differences in ecological validity. Further analysis reveals that ChatGPT's performance is highly context-dependent, shaped by specific application scenarios, model versions, and prompting strategies. To address methodological heterogeneity and the lack of standardization, this study recommends multi-method cross-validation strategies and a risk-stratified, standardized evaluation framework. These steps are essential to enhance the scientific rigor and reliability of ChatGPT's assessment in healthcare and to provide a solid foundation for its clinical integration.

Medical subject headings