Prompt Engineering and Follow-Up Questioning Improves the Readability of Spine Surgery Questions in Large Language Models.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 40889596.
- Also identified by DOI 10.1016/j.wneu.2025.124423.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
The field of spine surgery is complex, and patient education material within the field is often written at a reading level that exceeds the recommended standard. Large language models (LLMs), such as ChatGPT, have shown potential for generating educational content, but require further investigation to determine whether prompt-engineering or asking follow-up questions can improve readability. The purpose of this study is to evaluate which models, including newer versions, and prompting strategies have the greatest improvement in the readability. ChatGPT 4o and 5 were prompted with 45 standardized spine surgery questions across 5 common procedures. Each question underwent 5 prompting phases: baseline (phase 1), follow-up clarification (phase 1.5), sixth-grade level request (phase 2), rule-based prompting (phase 3), and direct readability targeting (phase 4). Readability was measured using Simple Measure of Gobbledygook, Flesch Reading Ease, Flesch-Kincaid Grade Level, Gunning Fog Index, and Coleman-Liau Index scoring. Results were standardized and analyzed using paired t-test and Wilcoxon signed-rank test, post-hoc analysis, of variance, and interaction models. Furthermore, a resource for spine surgeons with the most readable answers was created. While ChatGPT 4o generated significantly more readable responses than ChatGPT 5 across all phases except phase 1 (P < 0.001), ChatGPT 5 produced significantly more reliable citations (P < 0.05). Phase 2 yielded the most readable responses, with 51.11% meeting the sixth-grade level. Follow-up clarification questions and simplified prompt engineering were more effective than complex rule-based prompts. Prompt engineering and follow-up questioning significantly enhanced the readability of LLM-generated responses to spine surgery questions at a level appropriate for patients. Further investigation is needed to assess whether these responses are more readable in real-world clinical settings instead of relying on objective readability scoring methods.
Medical subject headings
- Comprehension
- Spine
- Language
- Patient Education as Topic
- Neurosurgical Procedures