Evaluating base and retrieval augmented LLMs with document or online support for evidence based neurology.
cross_sectional · Level IV
Where this comes from
- Record sourced from PubMed, PMID 40038423.
- Also identified by DOI 10.1038/s41746-025-01536-y and PMC identifier 11880332.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Effectively managing evidence-based information is increasingly challenging. This study tested large language models (LLMs), including document- and online-enabled retrieval-augmented generation (RAG) systems, using 13 recent neurology guidelines across 130 questions. Results showed substantial variability. RAG improved accuracy compared to base models but still produced potentially harmful answers. RAG-based systems performed worse on case-based than knowledge-based questions. Further refinement and improved regulation is needed for safe clinical integration of RAG-enhanced LLMs.