Embedding LLMs in the patient portal to summarize acute minor illness information: a three-arm experimental study.

Esmaeilzadeh, Pouyan · Int J Med Inform · 2026

rct · Level II

Where this comes from

Abstract

Patient portals provide direct access to clinical encounter data; however, comprehension barriers, especially for patients with limited health literacy, frequently prevent meaningful use, even for common acute conditions. Large language models (LLMs) embedded in the portal interface may convert clinical documentation into accessible plain-language summaries, yet no experimental study has evaluated this approach for acute minor illness. This study evaluated the impact of embedding two commercially available frontier LLMs (Claude Sonnet 4.5 [Anthropic] and GPT-5.1 Thinking [OpenAI]) into the patient portal of a small primary care clinic on patients' comprehension of information about acute minor illnesses, portal engagement, self-management adherence, and unnecessary healthcare utilization. We conducted a three-arm experimental study at a five-physician primary care practice. Adults (≥18 years) presenting with acute upper respiratory infection, influenza-like illness, or similar acute minor ailment were assigned 1:1:1 to: (A) Claude Sonnet 4.5 portal summaries, (B) GPT-5.1 Thinking portal summaries, or (C) standard portal access (control). The primary outcome was health information comprehension at Day 14, assessed via a study-specific, pilot-tested 10-item quiz (0-100 scale). A clinician audit of 20% of LLM summaries assessed accuracy and safety. Of 186 enrolled participants, 174 completed the 4-week follow-up (93.5% retention). At Day 14, comprehension scores were significantly higher in LLM arms versus control (Claude: M = 81.4, SD = 9.3; GPT-5.1: M = 79.6, SD = 10.1; control: M = 63.2, SD = 12.8; F(2,169) = 47.34, p < 0.001, partial η<sup>2</sup> = 0.36). Both LLM arms showed significantly higher total portal login frequency (Claude: p = 0.001; GPT: p = 0.007), longer session durations, and greater self-management adherence (p = 0.007) compared with the control. Unnecessary return visits were approximately 61% lower in combined LLM arms (OR = 0.34, 95% CI [0.14-0.81], p = 0.015). Minor factual imprecision rates in audited summaries were 5.7% (Claude) and 11.8% (GPT-5.1), with zero clinically significant hallucinations across all audited summaries. No significant difference was observed between the two LLMs. LLM-embedded patient portal summaries significantly improved comprehension of acute minor illnesses and reduced unnecessary healthcare utilization in a primary care setting. Because the intervention combined an LLM-generated plain-language summary with its prominent presentation in a dedicated portal panel, the observed effects should be interpreted as the result of this combined plain-language-plus-interface intervention rather than as a test of LLM capability alone. Both models performed comparably and safely, supporting the feasibility of further evaluation in diverse clinical settings. Further multi-site replication and implementation science studies are needed.