Evaluation of the performance of large language models in responding to medical questions related to multiple sclerosis: A case study of large language models including ChatGPT, Gemini, Grok and Copilot.

Dastani, Meisam; Sajjadi, Mohammad Shayan; Yamout, Bassem; Arab Bafrani, Melika; Nasirzadeh, Amirreza · PLoS One · 2026

other · Level V

Where this comes from

Abstract

To evaluate and compare the performance of four publicly available large language models-ChatGPT, Gemini, Copilot, and Grok-in answering medical questions related to Multiple Sclerosis, focusing on accuracy, transparency, and clinical actionability. Four publicly available large language models (ChatGPT, Gemini, Grok, and Copilot) were selected based on accessibility and their ability to respond to medical questions. A total of 25 questions-five for each of the five key domains (diagnosis, treatment, prevention, disease control, and disease management)-were developed. The responses generated by the models were evaluated using the DISCERN-AI and NLAT-AI assessment tools. The evaluation of four AI chatbots-ChatGPT, Gemini, Copilot, and Grok-on multiple sclerosis (MS) content revealed clear differences in quality and consistency. According to DISCERN-AI criteria, Gemini achieved the highest overall quality, excelling in relevance, transparency, balance, and acknowledgment of uncertainty. Grok ranked second, showing generally balanced results with slightly lower scores than Gemini. ChatGPT exhibited strong yet uneven performance, with particular weaknesses in content addressing vulnerable populations. Copilot demonstrated the weakest overall performance, with consistently lower scores across nearly all criteria. Gemini demonstrated the strongest and most consistent performance across all domains, followed by Grok with slightly lower but balanced results. ChatGPT showed strong yet uneven outcomes, with weaknesses in addressing vulnerable populations. Copilot ranked lowest, consistently underperforming across metrics. These findings highlight significant differences among large language models in generating accurate and clinically relevant responses for multiple sclerosis, underscoring the importance of considering each model's strengths and limitations in healthcare applications.

Medical subject headings