Transforming Neurosurgical Practice with Large Language Models: Comparative Performance of ChatGPT-Omni and Gemini in Complex Case Management.

Çöllüoğlu, Barış; Dikici, Şamil · World Neurosurg · 2025

prospective_cohort · Level II

Where this comes from

Abstract

Recent advancements in artificial intelligence, particularly in large language models (LLMs), have catalyzed new opportunities within medical domains, including neurosurgery. This study aims to evaluate and compare the performance of two advanced LLMs-ChatGPT-Omni and Gemini-in addressing clinical case inquiries on various neurosurgical conditions. A prospective observational study was conducted utilizing 500 case-based questions relevant to neurosurgery, covering 10 prevalent conditions. The questions were designed to simulate real-world clinical scenarios encompassing diagnosis, interpretation, and management and were asked again two months (phase 2) later. Responses were evaluated using a 6-point Likert scale by two independent neurosurgeons. ChatGPT-Omni exhibited consistent superiority across all evaluation metrics. In phase 1, its overall average score across all conditions was 5.38 ± 0.12, which increased to 5.46 ± 0.08 in phase 2 (P < 0.001). While exhibiting moderate improvements, Gemini trailed behind ChatGPT-Omni with an overall average score of 4.93 ± 0.15 in phase 1, which improved to 5.1 ± 0.14 in phase 2 (P < 0.001). Subgroup analyses indicated that ChatGPT-Omni provided superior contextual accuracy across all conditions (P < 0.001). The study underscores the transformative potential of LLMs in neurosurgery, with ChatGPT-Omni demonstrating superior accuracy, relevance, and clarity compared to Gemini. While both models improved over time, ChatGPT-Omni consistently excelled across all clinical scenarios, highlighting its potential utility in neurosurgical decision support and education.

Medical subject headings