Plain-language prompting improves readability of ChatGPT-generated patient education for meniscal surgery without loss of accuracy.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 42106008.
- Also identified by DOI 10.1016/j.jisako.2026.101132.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Large language models (LLMs) such as ChatGPT are increasingly used to generate patient education materials; however, default ChatGPT responses often exceed recommended readability levels set by the American Medical Association (AMA) and National Institutes of Health (NIH) health-literacy recommendations. The purpose of this study was to evaluate the readability and educational quality of ChatGPT-generated patient education on meniscal surgery and to determine whether a standardized plain-language prompt could improve readability without compromising accuracy, relevance, or depth. Sixteen standardized patient-focused questions regarding diagnosis, management, and prevention of meniscus tears were submitted to ChatGPT-5 and ChatGPT-4o, with three replicates per question to ensure standardization. Responses were assessed for accuracy against the American Academy of Orthopaedic Surgeons (AAOS) OrthoInfo and scored for relevance and depth using 5-point Likert scales. Readability was assessed using Flesch-Kincaid Grade Level (FKGL) and Flesch Reading Ease Score (FRES). All baseline responses were subsequently rewritten using a plain-language prompt targeting a sixth to eighth grade reading level. Pre- and post-prompt readability metrics were compared using paired t-tests. Inter-rater reliability was measured with Cohen's kappa. Both models demonstrated 100% factual accuracy across all baseline responses compared with OrthoInfo. Mean relevance and depth scores were high for ChatGPT-5 (4.49 ± 0.22; 4.39 ± 0.26) and ChatGPT-4o (4.55 ± 0.15; 4.53 ± 0.08). Baseline readability exceeded recommendations (FKGL 11.4-12.1; FRES 37-40). The plain-language prompt significantly improved readability for both models, reducing FKGL by approximately 5 grade levels and increasing FRES by 30-38 points (P < 0.001), with no loss of accuracy, relevance, or depth. ChatGPT generates accurate and relevant patient-directed content on meniscal surgery; however, readability frequently exceeded established health literacy standards. A simple, reproducible plain-language prompt reliably reduced reading level into the target range, offering a practical strategy for sports medicine surgeons to enhance informed consent discussions and patient education materials. IV.