A Comparison of Ten Large Language Models and a Conventional Search Engine for Clinical Decision Support in Anesthesiology: Expert Agreement and Physician Perceptions.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 41610406.
- Also identified by DOI 10.1213/ANE.0000000000007864.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Advances in artificial intelligence (AI) have enabled large language models (LLMs) to generate complex and contextually relevant medical responses. However, their potential in clinical decision support for anesthesiology remains underexplored. This study evaluated the accuracy and clinical relevance of high-performing LLMs in response to anesthesia-related questions and compared their performance with traditional online search methods. Clinician perceptions of AI were also assessed. We hypothesized that top-performing large language models would outperform lower-tier models and traditional internet search tools by generating responses rated as more accurate, complete, and clinically relevant to anesthesiology-focused questions, as measured by higher mean evaluator scores on a 10-point Likert scale. Ten LLMs: GPT-4o, Claude-Sonnet 3.5, DeepSeek R1, Llama 3.1 Instruct 70B, Gemini 2.0, GPT o1-preview, GPT o1, GPT o3-mini, NOVA Pro, and Mistral. All models were tested using ten common general anesthesia questions developed by TMH and validated by 6 physicians. Two Google search conditions served as baselines: a default search conducted in a cleared browser (unpersonalized), and a personalized Google Snippet Search performed in a browser regularly used by a clinician. Four board-certified anesthesiologists independently rated each response on a 10-point Likert scale. An ad hoc Physician Perception Questionnaire captured clinicians' use of AI, trust in its output, and reliance on traditional information sources. LLM performance varied significantly (F = 5.89, P < .0001). DeepSeek R1 achieved the highest overall score (7.7), whereas Gemini 2.0 Flash recorded the lowest among LLMs (5.2). The Google Snippet Search scored 5.3, the lowest overall. Pairwise Welch's t tests showed that DeepSeek R1 significantly outperformed Llama, o3-mini, and Mistral ( P < .001). Survey results indicated limited AI use in clinical practice; clinicians prioritized source credibility and continued to favor traditional resources. Although LLM-generated responses differed in quality, DeepSeek R1 and Claude-Sonnet 3.5 produced answers most consistent with expert clinical judgment. The poor performance of several models, coupled with clinician skepticism, underscores the need for further validation before integrating AI into routine anesthesiology decision support.
Medical subject headings
- Decision Support Systems, Clinical
- Anesthesiology
- Search Engine
- Anesthesiologists
- Artificial Intelligence
- Attitude of Health Personnel
- Physicians
- Language