Prompt Engineering and Follow-Up Questioning Improves the Readability of Spine Surgery Questions in Large Language Models.

Daulat, Sohail Rajesh; Dholaria, Nikhil; Burnet, Gregory; Patil, Shaunak; Manne, Bhavesh; Choudhary, Aditi; Mitha, Rida; Zeeshan, Qazi et al. · World Neurosurg · 2025

other · Level V

Where this comes from

Abstract

The field of spine surgery is complex, and patient education material within the field is often written at a reading level that exceeds the recommended standard. Large language models (LLMs), such as ChatGPT, have shown potential for generating educational content, but require further investigation to determine whether prompt-engineering or asking follow-up questions can improve readability. The purpose of this study is to evaluate which models, including newer versions, and prompting strategies have the greatest improvement in the readability. ChatGPT 4o and 5 were prompted with 45 standardized spine surgery questions across 5 common procedures. Each question underwent 5 prompting phases: baseline (phase 1), follow-up clarification (phase 1.5), sixth-grade level request (phase 2), rule-based prompting (phase 3), and direct readability targeting (phase 4). Readability was measured using Simple Measure of Gobbledygook, Flesch Reading Ease, Flesch-Kincaid Grade Level, Gunning Fog Index, and Coleman-Liau Index scoring. Results were standardized and analyzed using paired t-test and Wilcoxon signed-rank test, post-hoc analysis, of variance, and interaction models. Furthermore, a resource for spine surgeons with the most readable answers was created. While ChatGPT 4o generated significantly more readable responses than ChatGPT 5 across all phases except phase 1 (P < 0.001), ChatGPT 5 produced significantly more reliable citations (P < 0.05). Phase 2 yielded the most readable responses, with 51.11% meeting the sixth-grade level. Follow-up clarification questions and simplified prompt engineering were more effective than complex rule-based prompts. Prompt engineering and follow-up questioning significantly enhanced the readability of LLM-generated responses to spine surgery questions at a level appropriate for patients. Further investigation is needed to assess whether these responses are more readable in real-world clinical settings instead of relying on objective readability scoring methods.

Medical subject headings