Evaluation of the performance of large language models in responding to medical questions related to multiple sclerosis: A case study of large language models including ChatGPT, Gemini, Grok and Copilot.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 42113857.
- Also identified by DOI 10.1371/journal.pone.0346445 and PMC identifier 13160313.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
To evaluate and compare the performance of four publicly available large language models-ChatGPT, Gemini, Copilot, and Grok-in answering medical questions related to Multiple Sclerosis, focusing on accuracy, transparency, and clinical actionability. Four publicly available large language models (ChatGPT, Gemini, Grok, and Copilot) were selected based on accessibility and their ability to respond to medical questions. A total of 25 questions-five for each of the five key domains (diagnosis, treatment, prevention, disease control, and disease management)-were developed. The responses generated by the models were evaluated using the DISCERN-AI and NLAT-AI assessment tools. The evaluation of four AI chatbots-ChatGPT, Gemini, Copilot, and Grok-on multiple sclerosis (MS) content revealed clear differences in quality and consistency. According to DISCERN-AI criteria, Gemini achieved the highest overall quality, excelling in relevance, transparency, balance, and acknowledgment of uncertainty. Grok ranked second, showing generally balanced results with slightly lower scores than Gemini. ChatGPT exhibited strong yet uneven performance, with particular weaknesses in content addressing vulnerable populations. Copilot demonstrated the weakest overall performance, with consistently lower scores across nearly all criteria. Gemini demonstrated the strongest and most consistent performance across all domains, followed by Grok with slightly lower but balanced results. ChatGPT showed strong yet uneven outcomes, with weaknesses in addressing vulnerable populations. Copilot ranked lowest, consistently underperforming across metrics. These findings highlight significant differences among large language models in generating accurate and clinically relevant responses for multiple sclerosis, underscoring the importance of considering each model's strengths and limitations in healthcare applications.
Medical subject headings
- Multiple Sclerosis
- Large Language Models