From Guidelines to Clicklists: GPT-5-Generated ERAS Checklists Improve Guideline Coverage for Bariatric and Gastrointestinal Cancer Surgery-A STROBE-Compatible Cross-Sectional Evaluation.
cross_sectional · Level IV
Where this comes from
- Record sourced from PubMed, PMID 41873099.
- Also identified by DOI 10.1002/wjs.70339.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Enhanced Recovery After Surgery (ERAS) pathways improve outcomes after bariatric and gastrointestinal (GI) cancer surgery, yet real-world adherence remains inconsistent. Digital tools and checklists can support implementation, but their maintenance and completeness may be limited. Large language models (LLMs) could rapidly generate structured ERAS checklists; however, their coverage, clarity, and bias profile require systematic evaluation. A practical concern is "bundle inflation," whereby expanding item counts may undermine feasibility even when individual elements are evidence-based. We performed a STROBE-compatible cross-sectional observational study (March-June 2025) evaluating AI-generated ERAS checklists against guideline-derived comparators. Using GPT-5, we generated 12 ERAS checklists (6 bariatric; 6 GI cancer: 3 gastrectomy and 3 colorectal). Twelve traditional checklists were curated from ERAS Society guideline items. Three blinded raters (two board-certified surgeons; one clinical informatics specialist) independently scored item coverage (present/absent per guideline item), clarity (5-point Likert), and potential bias/applicability issues using a predefined rubric. Primary outcomes were guideline-item coverage (%) and clarity. Interrater reliability was assessed using Cohen's kappa; group comparisons used two-sided tests with α = 0.05. Coverage was not weighted; all guideline items contributed equally, and "critical" omissions were defined a priori as omissions of items explicitly labeled "strong" or "recommended" in the source guideline documents. AI-generated checklists demonstrated higher mean guideline-item coverage than traditional checklists (97.0% ± 2.1% vs. 89.0% ± 3.2%) and higher clarity scores (4.8 ± 0.2 vs. 4.2 ± 0.3; p = 0.021). Agreement was excellent (κ = 0.92; 95% CI 0.88-0.97). Raters observed no systematic demographic bias; limitations primarily reflected reduced context-specific tailoring (e.g., nutrition pathways for selected subgroups). AI-generated lists contained slightly more discrete items than traditional templates, highlighting the need for an implementability review to prevent overly long bundles from reducing adherence. GPT-5-generated ERAS checklists achieved superior guideline coverage and clarity versus traditional checklists in bariatric and GI cancer surgery. Because more items can paradoxically reduce implementation fidelity, AI outputs should be treated as draft "master lists" that require structured local curation (core vs. conditional elements) before deployment. Prospective, workflow-integrated validation is warranted to confirm real-world effectiveness and mitigate context-specific omissions.