Large language models for interpretation of health checkup results.

You, Jiwon; Shin, Hangsik · NPJ Digit Med · 2026

basic_science · Level V

Where this comes from

Abstract

Large language models (LLMs) show strong generalization, yet their ability to interpret structured medical data remains insufficiently studied. This work evaluated four LLMs-Claude Sonnet 4, Gemini 2.5 Pro, GPT-4o, and LLaMA 3.1-70B-using comprehensive health checkup data from the Korean National Health Insurance Service. Multiple prompting strategies (few-shot, role-based, constraint-based, and Chain-of-Thought) were tested. Zero-shot accuracy averaged 0.69 (SD 0.06), increasing to 0.92 (0.06) with combined strategies and to 0.95 (0.07) with Chain-of-Thought. Claude Sonnet 4, Gemini 2.5 Pro, and GPT-4o achieved the highest accuracies (≥ 0.98), while LLaMA 3.1-70B showed lower but improvable performance. Item-level analysis of 10,000 cases demonstrated near-perfect accuracy (0.99-1.00) for most biochemical markers, including glucose, cholesterol, triglycerides, and liver enzymes. In contrast, blood pressure showed lower accuracy (0.61-0.91), with age-related decline, likely due to the complexity of multi-categorical thresholds requiring integration of systolic and diastolic values. Subgroup analyses revealed model-specific biases: sex-related biases were observed in body mass index (Claude Sonnet 4) and urine protein, serum creatinine, and gamma-glutamyl transferase (LLaMA 3.1-70B), while age-related biases were identified in blood pressure (Claude Sonnet 4, Gemini 2.5 Pro) and low-density lipoprotein cholesterol and hemoglobin (LLaMA 3.1-70B). Overall, advanced prompting markedly improved model performance, and top-tier LLMs demonstrated robust interpretive capability, while caution is needed for variables with complex clinical semantics such as blood pressure.