Too agreeable to be accurate? Sycophancy and diagnostic instability of large language models in medical diagnosis.

Samsel, Konrad; Dharma, Christoffer; Rezaei, AmirHossein H M; Bahakim, Aseel; Ravikumar, Kynthia; Bhat, Venkat; Kamaleddinn, Muhammad Amin; Shakeri, Zahra · Artif Intell Med · 2026

other

Where this comes from

Abstract

Diagnostic accuracy in Large Language Models (LLM) is an increasing concern as physicians employ LLMs into their medical practice. We measured whether diagnostic correctness in LLMs changed after a certainty challenge or clinician specialty framing. To evaluate the performance of LLMs, we curated 120 public clinical vignettes: 40 MultiCaRe-derived clinical narratives and 80 MedMCQA-derived exam style cases. Ten proprietary and open-weight LLMs were tested across four prompt conditions and three separate runs with two passes per condition, yielding 28,800 responses. Diagnostic correctness was evaluated using an LLM-as-a-judge approach. At neutral baseline, accuracy varied by model and was generally lower for MultiCaRe than MedMCQA. GPT-5 had the highest baseline and post-challenge accuracy (74.4% and 75.3%); Claude Sonnet 4 had the highest accuracy changing flip rate (AcFR; 58.6%) and largest post-challenge accuracy loss (-36.4 percentage points). Across all neutral model-case pairs, baseline accuracy decreased from 51.8% to 42.2% after 'Are you sure?' (AcFR = 32.8%). Correct-to-incorrect transitions (n = 544) exceeded incorrect-to-correct transitions (n = 198). Specialty context prompts produced smaller changes. Adjacent specialty prompts increased overall accuracy by 2.8 percentage points, and differential specialty prompts decreased accuracy by 1.2-1.8 percentage points. The 'Are you sure?' challenge was the strongest source of diagnostic instability in the study; case source also strongly affected accuracy. Diagnostic LLM evaluations should report first pass accuracy, harmful flips, beneficial corrections, and stability under clinically plausible conversational challenges before clinical deployment.