Building trustworthy large language model-driven generative recommender system for healthcare decision support: A scoping review of corpus sources, customization techniques, and evaluation frameworks.

Yang, Shuqi; Jing, Mingrui; Wang, Shuai; Huang, Zongan; Wang, Jiaqing; Kou, Jiaxin; Shi, Manfei; Xia, Zhentao et al. · Artif Intell Med · 2026

review · Level V

Where this comes from

Abstract

Large Language Model-Driven Generative Recommender Systems (LLM-GRSs) are playing a growing role in healthcare, particularly in clinical question-answering. This study reviews their corpus sources, customization techniques, and evaluation metrics. We conducted a systematic search of PubMed, Embase, Scopus, and Web of Science for studies published between January 2021 and August 2025 that applied LLM-GRSs to deliver medical or healthcare information. Eligible studies included publications describing LLMs designed to emulate clinical decision-making by providing diagnostic or therapeutic recommendations through dialogue-based interfaces. Two reviewers independently screened studies and extracted data on corpus sources, model architectures, customization methods, and evaluation metrics. A total of 61 articles were included. Corpus sources were grouped into clinical data (n = 25), literature (n = 34), open datasets (n = 37), and web-crawled data (n = 15), with many using multiple types. Most studies (n = 43) combined multiple approaches. Customization techniques included prompt engineering, retrieval-augmented generation and model fine-tuning. Twenty-four studies used a single customization technique, while 37 studies combined these methods during model development. The evaluation metrics were classified into three main domains: process metrics, usability metrics, and outcome metrics. The outcome metrics included both model-based and manual-assessed evaluations. LLM-GRSs hold considerable promise in healthcare; however, their safety and reliability hinge on the use of evidence-based training corpora, transparent system design, and standardized evaluation protocols within real-world clinical environments.

Medical subject headings