Enhancing AI for citation screening in literature reviews: Improving accuracy with ensemble models.

Zhang, Zhihong; Momeni Nezhad, Mohamad Javad; Gupta, Pallavi; Zolnour, Ali; Azadmaleki, Hossein; Topaz, Maxim; Zolnoori, Maryam · Int J Med Inform · 2025

basic_science · Level V

Where this comes from

Abstract

Healthcare literature reviews underpin evidence-based practice and clinical guideline development, with citation screening as a critical yet time-consuming step. This study evaluates the effectiveness of individual large language models (LLMs) versus ensemble approaches in automating citation screening to improve the efficiency and scalability of evidence synthesis in healthcare research. Performance was assessed across three healthcare-focused reviews: LLM-Healthcare (865 citations, broad scope, 49.8 % inclusion rate), MCI-Speech (959 citations, narrow scope, 6.5 % inclusion rate), and Multimodal-LLM (73 citations, moderate scope, 68.5 % inclusion rate). Six LLMs (GPT-4o Mini, GPT-4o, Gemini Flash, Llama 3.1 8B Instruct, Llama 3.1 70B Instruct, Llama 3.1 405B Instruct) were evaluated using zero- and few-shot learning strategies with PubMedBERT for demonstration selection. We compared individual model performance with ensemble methods, including majority voting and random forest (RF), based on sensitivity and specificity. No individual LLM consistently outperformed others across all tasks. Review with narrow inclusion criteria and low inclusion rates exhibited high specificity but lower sensitivity. Ensemble methods consistently surpassed individual LLMs: the RF ensemble with GPT-4o performed best in LLM-Healthcare (sensitivity: 0.96, specificity: 0.89); the majority voting with 1-shot LLMs (sensitivity: 0.75, specificity: 0.86) and RF ensemble with 4-shot LLMs (sensitivity: 0.62, specificity: 0.97) excelled in MCI-Speech; and four RF ensembles achieved perfect classification (sensitivity: 1.0, specificity: 1.0) in Multimodal-LLM. Ensemble approaches improve individual LLMs' performances in citation screening across diverse healthcare review tasks, highlighting their potential to enhance evidence synthesis workflows that support clinical decision-making. However, broader validation is needed before real-world implementation.

Medical subject headings