Cross-model consistency of AI-generated exercise prescriptions: A repeated generation study across three large language models.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 42435611.
- Also identified by DOI 10.1016/j.ijmedinf.2026.106594.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Large language models (LLMs) are increasingly applied to exercise prescription, yet cross-model differences in output consistency under repeated generation conditions remain unexamined. This study aimed to systematically compare repeated generation consistency of exercise prescription outputs across three widely used LLMs under identical conditions. GPT-4.1, Claude Sonnet 4.6, and Gemini 2.5 Flash each generated prescriptions for six clinical scenarios 20 times (360 total outputs) under temperature = 0 conditions. Outputs were analyzed across four dimensions: semantic similarity (SBERT cosine similarity), output reproducibility, FITT component classification, and safety expression. Mean semantic similarity was highest for GPT-4.1 (0.955), followed by Gemini 2.5 Flash (0.950) and Claude Sonnet 4.6 (0.903), with significant inter-model differences confirmed (H = 458.41, p < 0.001, ε<sup>2</sup> = 0.134). These scores reflected fundamentally different generative behaviors: GPT-4.1 produced entirely unique outputs (100%) with stable semantic content, while Gemini 2.5 Flash showed pronounced output repetition (27.5% unique outputs), indicating that its high similarity score derived from text duplication rather than consistent reasoning. Safety expression reached ceiling levels across all models (mean Safety Total: 3.93-3.99 out of 4.00), confirming its limited utility as a differentiating metric. These findings suggest that model-specific output behavior should be considered when evaluating LLMs for exercise prescription support, particularly in applications requiring reproducible and guideline-compatible outputs.