Evaluating a Guideline-Integrated Clinical Interaction Framework Vs a Standard Large Language Model Interaction for Dietary Recommendations in Recurrent Urolithiasis: In Silico Study.

Wang, Xiaofeng; Li, Jun; Hu, Yudong; Chen, Yujie; Zhong, Yong; Zhu, Faming; Yuan, Ye; Yang, Fan et al. · J Med Internet Res · 2026

other · Level V

Where this comes from

Abstract

Personalized dietary counseling is central to recurrence prevention in patients with urolithiasis, particularly after a 24-hour urine metabolic evaluation. However, translating quantitative metabolic abnormalities into patient-facing, guideline-concordant, and safe dietary recommendations can be challenging in routine clinical practice. Large language models (LLMs) may assist with this task, but unguided responses may overlook key metabolic priorities or case-specific safety constraints. This study evaluated whether a guideline-integrated, safety-aware, LLM-based clinical interaction framework (StoneAgent) could generate higher-quality, personalized dietary recommendations than a standard LLM configuration for recurrent urolithiasis. We also assessed whether any performance advantage persisted when the same clinical scenarios were presented as patient query-style inputs. We conducted an in silico comparative study using 30 synthetic clinical vignettes representing common, mixed, and safety-relevant metabolic stone scenarios. For the primary experiment, StoneAgent and a standard LLM configuration were compared using structured vignette inputs. For the robustness experiment, each vignette was reformulated into 2 patient query-style variants (query A and query B), preserving the same clinical content in more natural conversational language. A vignette-specific expert reference standard was developed from guideline-informed specialist consensus. Three independent reviewers blindly rated outputs on a 5-point Likert scale for metabolic specificity, guideline adherence, and actionability; safety was assessed as a binary outcome. For the patient query-style experiment, query A and query B were aggregated at the vignette level for paired comparison. In the structured-input experiment, StoneAgent achieved higher performance than the standard LLM across metabolic specificity, guideline adherence, and actionability, with median case-level scores of 5.00 (IQR 5.00-5.00) vs 3.00 (IQR 2.75-3.92) for metabolic specificity, 5.00 (IQR 5.00-5.00) vs 3.67 (IQR 3.08-4.00) for guideline adherence, and 5.00 (IQR 5.00-5.00) vs 3.00 (IQR 2.67-3.33) for actionability (all <i>P</i><.001). Safety pass rates were 100% (30/30) for StoneAgent and 83.3% (25/30) for the standard LLM (exact McNemar <i>P</i>=.06). In the patient query-style robustness experiment, StoneAgent retained a directional advantage after case-level aggregation, with mean scores of 4.44 vs 3.62 for metabolic specificity, 4.51 vs 3.63 for guideline adherence, and 4.11 vs 3.40 for actionability. Safety pass rates were 100% (30/30) for StoneAgent and 90% (27/30) for the standard LLM (exact McNemar <i>P</i>=.25). The performance gap was more conservative under patient query-style inputs compared with structured inputs, but the overall pattern remained consistent across domains. In this in silico study, a guideline-integrated, safety-aware clinical interaction framework generated higher-quality dietary recommendations for recurrent urolithiasis than a standard LLM condition with structured vignette inputs. This advantage was retained with patient query-style inputs. These findings suggest that explicit clinical framing, guideline grounding, and safety-oriented response scaffolding may improve the reliability of specialty counseling tasks involving metabolic stone prevention. Further validation is needed using real patient-authored queries and prospective clinical workflows.

Medical subject headings