Plain-language prompting improves readability of ChatGPT-generated patient education for meniscal surgery without loss of accuracy.

Chakraborty, Aritra; Mattin, Drew; Christiansen, Mitchell J; Mulcahey, Mary K · J ISAKOS · 2026

other · Level V

Where this comes from

Abstract

Large language models (LLMs) such as ChatGPT are increasingly used to generate patient education materials; however, default ChatGPT responses often exceed recommended readability levels set by the American Medical Association (AMA) and National Institutes of Health (NIH) health-literacy recommendations. The purpose of this study was to evaluate the readability and educational quality of ChatGPT-generated patient education on meniscal surgery and to determine whether a standardized plain-language prompt could improve readability without compromising accuracy, relevance, or depth. Sixteen standardized patient-focused questions regarding diagnosis, management, and prevention of meniscus tears were submitted to ChatGPT-5 and ChatGPT-4o, with three replicates per question to ensure standardization. Responses were assessed for accuracy against the American Academy of Orthopaedic Surgeons (AAOS) OrthoInfo and scored for relevance and depth using 5-point Likert scales. Readability was assessed using Flesch-Kincaid Grade Level (FKGL) and Flesch Reading Ease Score (FRES). All baseline responses were subsequently rewritten using a plain-language prompt targeting a sixth to eighth grade reading level. Pre- and post-prompt readability metrics were compared using paired t-tests. Inter-rater reliability was measured with Cohen's kappa. Both models demonstrated 100% factual accuracy across all baseline responses compared with OrthoInfo. Mean relevance and depth scores were high for ChatGPT-5 (4.49 ± 0.22; 4.39 ± 0.26) and ChatGPT-4o (4.55 ± 0.15; 4.53 ± 0.08). Baseline readability exceeded recommendations (FKGL 11.4-12.1; FRES 37-40). The plain-language prompt significantly improved readability for both models, reducing FKGL by approximately 5 grade levels and increasing FRES by 30-38 points (P < 0.001), with no loss of accuracy, relevance, or depth. ChatGPT generates accurate and relevant patient-directed content on meniscal surgery; however, readability frequently exceeded established health literacy standards. A simple, reproducible plain-language prompt reliably reduced reading level into the target range, offering a practical strategy for sports medicine surgeons to enhance informed consent discussions and patient education materials. IV.