Retrieval-augmented generation in medicine: A scoping review of technical implementations, clinical applications, and ethical considerations.

Yang, Rui; Wong, Matthew Yu Heng; Li, Huitao; Li, Xin; Zhu, Wentao; Liao, Jingchi; Yu, Kunyu; Liew, Jonathan Chong Kai et al. · Cell Rep Med · 2026

review · Level V

Where this comes from

Abstract

The rapid growth of medical knowledge and the increasing complexity of clinical practice pose challenges. In this context, large language models (LLMs) demonstrate value; however, inherent limitations remain. Retrieval-augmented generation (RAG) shows potential to enhance their clinical applicability. This study reviews RAG applications in medicine. We find that research primarily relies on publicly available data, with limited use of private data. For retrieval, approaches commonly rely on English-centric embedding models, while LLMs are mostly generic, with limited use of medical-specific LLMs. For evaluation, automated metrics evaluate generation quality and task performance, whereas human evaluation focuses on accuracy, completeness, relevance, and fluency, with insufficient attention to bias and safety. RAG applications are concentrated on question answering, report generation, text summarization, and information extraction. Overall, medical RAG remains at an early stage, requiring advances in clinical validation, cross-linguistic adaptation, and support for low-resource settings to enable trustworthy and responsible global use.