Benchmarking retrieval-augmented large language models in biomedical NLP: Application, robustness, and self-awareness.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41270181.
- Also identified by DOI 10.1126/sciadv.adr1443 and PMC identifier 12637297.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
To reduce hallucinations in large language models (LLMs), retrieval-augmented LLMs (RALs) retrieve supporting knowledge from external databases. However, their performance on biomedical natural language processing (NLP) tasks remains underexplored. We introduce Biomedical Retrieval-Augmented Generation Benchmark, a comprehensive evaluation framework assessing RALs across five biomedical NLP tasks and 11 datasets, using four testbeds: unlabeled robustness, counterfactual robustness, diverse robustness, and self-awareness. To improve RALs' robustness and negative awareness, we propose a detect-and-correct strategy and a contrastive learning approach. Experimental results show that RALs generally outperform standard LLMs on most biomedical tasks, but still struggle with robustness and self-awareness, particularly under counterfactual and diverse scenarios. Our proposed methods significantly improve performance in robustness to unlabeled and counterfactual data, and increase the model's ability to detect and avoid incorrect predictions. These findings highlight key limitations in current RALs and underscore the need for continued refinement to ensure reliability and accuracy in high-stakes biomedical applications.
Medical subject headings
- Natural Language Processing
- Benchmarking
- Information Storage and Retrieval