Fine-Tuning, Retrieval-Augmented Generation, and Hybrid Large Language Models for Postoperative Decision Support: A Comparative Analysis.

Prabha, Srinivasagam; Collaco, Bernardo Gabriele; Gomez-Cabello, Cesar Abraham; Haider, Syed Ali; Genovese, Ariana; Fang, Zhihui; Wood, Nadia; Bagaria, Sanjay et al. · J Med Internet Res · 2026

other

Where this comes from

Abstract

Large language models (LLMs) show growing potential for decision support. However, integrating domain-specific medical knowledge while maintaining accuracy, safety, and interpretability remains challenging for postoperative discharge instructions and patient education. Fine-tuning (FT), retrieval-augmented generation (RAG), and hybrid FT+RAG approaches are prominent strategies for knowledge integration, but their comparative performance in postoperative care has not been systematically evaluated. We aimed to compare the performance, reliability, and safety characteristics of baseline, FT, RAG, and hybrid FT+RAG LLM configurations for postoperative decision support. We conducted a comparative evaluation of four LLM configurations using Google Gemini 2.5 Flash. A total of 600 postoperative question-answer pairs were used for model adaptation and validation, while 150 queries were reserved for final evaluation. Queries included routine postoperative questions, emergency escalation scenarios, and deliberately out-of-scope prompts. Outputs were independently assessed by 3 blinded clinical experts for clinical medical accuracy, safety/refusal accuracy, completeness, and relevance. Automated metrics evaluated readability, faithfulness, and hallucination propensity. All knowledge-enhanced models significantly outperformed baseline in overall accuracy (baseline 68.0% vs FT 92.7%, RAG 91.3%, FT+RAG 97.3%; P<.001). For in-scope clinical queries, FT+RAG achieved the highest clinical medical accuracy (96.7%) and was the only configuration to significantly outperform baseline in pairwise comparisons. Enhanced models also demonstrated higher safety/refusal accuracy than baseline; however, the baseline configuration did not receive equivalent safety or deferral instructions, which likely influenced these findings. FT+RAG achieved the strongest composite classification performance, including 100% precision, 96.7% recall, and 98.3% F1 score. FT and RAG showed broadly comparable performance across most secondary outcomes. Although knowledge-enhanced models demonstrated lower readability than baseline, restricted analysis of 100 routine in-scope postoperative queries suggested that part of this difference was attributable to standardized safety boilerplate. Incorporating domain-specific knowledge through FT, RAG, or both improved postoperative decision-support performance compared with the baseline LLM. All knowledge-enhanced approaches demonstrated strong performance, with the hybrid FT+RAG configuration achieving the most favorable overall point estimates across several outcomes. However, differences among the enhanced configurations were generally modest and less evident in sensitivity analyses restricted to unanimously rated queries. These findings support knowledge-enhanced LLMs as promising tools for postoperative education and decision support, while highlighting the need for further validation, readability optimization, transparent governance, and sustained human oversight before patient-facing deployment.