Assessing the safety of patient-centred discharge medication instructions generated by an AI model.

Tang, Michael; Markova, Kristina; Stanceski, Kristian; Zhong, Sharleen; Tracy, Marguerite; Koria, Linda; Lo, Sarita; Zhang, Xumou et al. · Int J Med Inform · 2026

cross_sectional · Level IV

Where this comes from

Abstract

The use of AI to support patient-centred communication could improve health outcomes but little is known about the equity of AI tools. We evaluated the completeness and accuracy of an AI tool that produces patient-centred medication information for patients following discharge from hospital, for different patient groups. We evaluated differences in the completeness and safety of AI-generated (GPT-4o) patient-centred medication instructions across age groups, patient complexity, and insurance type. AI-generated medication instructions were evaluated by clinical experts for the proportion of medications that were correctly represented, described in Universal Medication Schedule (UMS) form, and presence of safety issues. We tested for significant differences in completeness and safety between groups in 140 discharge summaries sampled from the Medical Information Mark for Intensive Care (MIMIC) database. The proportion of patient-centred discharge instructions where all medications were included was 95 % (133/140) with a median of 6.0 medications (IQR 3.0-10.0). For most of the 140 cases, all medications from the discharge summary were correctly included (median 100 % included, IQR 83.3 %-100 %) and new medications were rarely added by AI, but a lower proportion of medications were presented in UMS format (median 22.5 %, IQR 0.0 %-92.5 %). Despite most medications being included, potential safety issues were identified in 69.3 % (97/140). There was no evidence of a difference in the correctness of included medications across age groups (p = 0.70), patient complexity (p = 0.72), or insurance type (p = 0.70). There was no evidence of a difference in proportion of medications in UMS format across age groups (p = 0.88), patient complexity (p = 0.94), or insurance type (p = 0.49). There was evidence of a difference in the proportion of cases with at least one potential safety issue across age groups (p = 0.031), patient complexity (p < 0.001) and insurance types (p = 0.047). We found evidence of a difference in safety issues in AI-generated medication instructions for older, more complex patients, and patients with certain types of insurance. Health system and contextual differences could create unexpected variations in AI-generated outputs. Studies of AI-generated messaging for patients should consider the severity and likelihood of safety issues, localised trials, and ongoing auditing.

Medical subject headings