Clinical agents fail silently on patient identity.

Klang, Eyal; Glicksberg, Benjamin S; Gorenshtein, Alon; Gavin, Nicholas; Freeman, Robert; Stump, Lisa; Charney, Alexander W; Wei Ting, Daniel Shu et al. · Int J Med Inform · 2026

other · Level V

Where this comes from

Abstract

LLMs now power clinical agents that plan, call tools, and write into EHRs in routine clinical workflows today. Before clinical deployment, we must know whether agents detect patient identity faults or write through them. We built a record environment from publicly available MIMIC-IV emergency department data. Agents copied ICD-10-CM codes from visit headers into patient records using Extract and Store tools, with "UNKNOWN" allowed when uncertain. Missing Store calls were counted as non-writes. Each run presented ten visits from one patient (clean), then one visit was tampered. We tested four conditions: clean baseline, one visit with a fully swapped header from another patient, one visit with a one-digit MRN change, and one visit with age shifted. Six models, spanning closed and open weights, completed 1.2 million tool calls. Tamper detection was defined as withholding a write on the tampered visit by outputting "UNKNOWN" or omitting the Store call. Agents usually copied codes into tampered charts. On the tampered visit, GPT-4.1 withheld writing in 17.4% of header-swap runs, but detection of subtle faults (MRN or age changes) was near zero. GPT-4.1-nano detected 4.4% of header swaps and < 1% of MRN or age changes. GPT-5-chat never identified mismatches but produced non-writes in 12.6% of cases. Other models rarely withheld writing. Clinical agents often fail to detect patient identity inconsistencies. The central risk is misbinding, not miscoding. Safe deployment requires explicit identity verification, abstention when uncertain, and benchmarks that treat record integrity, not just accuracy, as a primary outcome.