From GPT-4 to Expert-Endorsed Athlete Guidance: A Delphi Consensus on Sleep and Jet Lag.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 42470603.
- Also identified by DOI 10.1007/s40279-026-02484-7.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Large language models (LLMs) including GPT-4 are increasingly used to generate health information, but concerns persist about their accuracy and relevance, particularly for elite athletes. This study used GPT-4-generated frequently asked question (FAQ) responses on sleep and jet lag as the starting material for expert evaluation and consensus development, with the goal of producing consensus-based, athlete-specific guidance while identifying the limitations of AI-generated content. Between November 2024 and March 2025, n = 17 international sleep and circadian experts from the Athlete Travel & Sleep Interest Group (ATSIG) participated in a two-round Delphi process. Experts rated 20 GPT-4-generated FAQ responses (10 on sleep, 10 on jet lag) for appropriateness using a 6-point Likert scale and provided qualitative feedback. Items were revised after round 1 using inductive thematic coding. Consensus was defined as ≥ 70% of participants rating an item as appropriate (scores 5-6) and ≤ 15% as non-appropriate (scores 1-2). Statistical analyses included Wilcoxon signed-rank tests, convergence metrics and dissent detection (outlier and bipolarity analysis). In round 1, 15 of 20 items (75%) met consensus; by round 2, 18 of 20 (90%) achieved consensus. For sleep items, 7 of 10 reached consensus in round 1 and 9 in round 2; for jet lag, 8 items reached consensus in round 1 and 9 in round 2. Sleep Q6 (sleep and injury risk) narrowly missed the consensus threshold with 64.7% agreement, while Jet Lag Q9 (melatonin and sleep aids) remained below the 70% threshold. No item showed bimodal score distributions, suggesting no polarized disagreement. Descriptive rating patterns, increased consensus and qualitative expert feedback indicated improved clarity, accuracy and athlete-specific relevance after the Delphi process, although item-level statistical comparisons did not remain significant after Bonferroni correction for multiple testing. Qualitative analysis identified common concerns: for sleep items-imprecise or misleading content (55%), lack of athlete-specific relevance (30%) and outdated evidence (11%); for jet lag-outdated evidence (36%), imprecise or misleading content (34%) and formatting issues (17%). GPT-4-derived content may serve as useful preliminary material for expert discussion but should not be used as standalone guidance. Expert evaluation improved the clarity, safety and athlete-specific relevance of most sleep and jet-lag responses, while the final outputs should be interpreted as consensus-based guidance rather than definitive proof of correctness.