Real-world evaluation of medication recommendation workflows: Retrieval augmentation, physician-RAG collaborative workflow, and prescribing quality.
retrospective_cohort · Level III
Where this comes from
- Record sourced from PubMed, PMID 42456600.
- Also identified by DOI 10.1016/j.ijmedinf.2026.106598.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Large language models (LLMs) are increasingly explored for clinical decision support, yet their role in medication recommendation remains uncertain. Prescribing requires not only selecting appropriate therapies, but also prioritizing essential treatments, avoiding unsafe medications, and specifying correct prescription details. Existing evaluations often emphasize benchmark reasoning rather than prescribing readiness in real-world settings. To compare four medication recommendation workflows using a clinically grounded framework focused on completeness, safety, ranking quality, and prescription-parameter accuracy. We conducted a retrospective real-world evaluation of 800 clinical cases from Shanghai East Hospital. A case-specific expert reference standard was constructed using three mutually exclusive categories: CORE, essential therapies; ALT, acceptable alternatives; and AVOID, unsafe or inappropriate medications. Four workflows were compared: Physician, Base LLM, Retrieval-augmented LLM, and Physician-RAG collaborative. Recommendations were normalized at the drug level and evaluated using recall, case-level AVOID-hit rate, extra-drug burden, precision, Jaccard index, hit@k, and prescription-parameter accuracy. Performance differed across workflows. Physician-RAG collaborative showed the highest CORE recall (1.000; 95 % CI, 0.999-1.000), ALT recall (0.955; 95 % CI, 0.945-0.964), overall recall (0.978; 95 % CI, 0.973-0.982), precision (0.963; 95 % CI, 0.957-0.969), and Jaccard index (0.943; 95 % CI, 0.935-0.949). It also had the lowest case-level AVOID-hit rate (0.026; 95 % CI, 0.016-0.039) and the lowest ranked unsafe exposure (AVOID hit@1/3/5: 0.000/0.000/0.008). Prescription-parameter accuracy was highest in the Physician-RAG collaborative workflow (all-correct rate 0.951), compared with 0.948 for the Retrieval-augmented LLM workflow, 0.945 for the Physician workflow, and 0.494 for the Base LLM workflow. In this retrospective expert-reference-based evaluation, stand-alone LLM output showed lower agreement and weaker structured prescribing accuracy than supervised workflows. Retrieval augmentation improved coverage and prescribing accuracy, while Physician-RAG collaborative showed the highest agreement with the expert reference standard across completeness, safety screening, prioritization, and structured prescribing quality. These findings support physician-supervised collaboration as a promising direction for medication decision support, but prospective multicenter studies with patient-outcome endpoints are needed before routine clinical deployment.