Artificial Intelligence Simplification of English and Spanish Surgical Consent Forms: 1-Size Does Not Fit All.

Pettigrew, Morgan F; Nunez-Rocha, Ricardo E; Govindu, Sai R; Castillo, Samy; Heslin, Ryan T; Bain, Andrew P; Abreu, Andres A; Fatimah, Nafeesah et al. · J Am Coll Surg · 2026

other · Level V

Where this comes from

Abstract

Guidelines recommend patient-facing material be written at or below a 6th- to 9th-grade reading level; however, surgical consent forms frequently exceed this standard. Artificial intelligence chatbots offer accessible tools to enhance readability, but their performance for non-English material remains unclear. We evaluated 3 strategies to improve consent readability. First, ChatGPT-4.0 was prompted to simplify preexisting surgical consent forms in English and Spanish. Second, a custom generative pretrained transformer (GPT) was trained using these outputs and tested on a separate validation cohort. Finally, both GPT's were prompted to translate English forms into Spanish, targeting a 6th-grade reading level. Readability was measured by the Fry Readability Formula, Simple Measure of Gobbledygook (SMOG) Index, Flesch-Kincaid Grade Level for English, and the Gilliam-Peña-Mountain (GPM) and SOL Readability Formulas for Spanish. ChatGPT-4.0 modestly improved English readability, with median Fry decreasing from 12.0 to 10.0, SMOG 12.3 to 11.7, and Flesch-Kincaid 9.2 to 8.6 (all p < 0.05). Spanish readability did not significantly improve (median GPM 9.5 to 10.0; SOL 9.6 to 9.1; both p > 0.05).Our custom GPT improved English readability, reducing median Fry from 12.0 to 7.0, SMOG 12.3 to 9.3, and Flesch-Kincaid 9.6 to 5.9 (all p < 0.001). Spanish readability gains were modest with median SOL decreasing from 9.5 to 7.3 (p < 0.01). GPM was unchanged (8.0, p > 0.99). English-to-Spanish translation outperformed direct Spanish simplification. ChatGPT-4.0 achieved median GPM of 8.0 (p = 0.25) and SOL of 7.6 (p < 0.001, compared with original Spanish consents). Translation via the custom GPT achieved the largest improvement (median GPM 7.0 and SOL 6.8; both p < 0.05, compared with original Spanish consents). Tailored artificial intelligence chatbots can improve readability of patient-facing material; however, diverse strategies are necessary to adapt existing models for multilingual populations.

Medical subject headings