Agentic memory-augmented retrieval and evidence grounding for medical question-answering tasks.

Jia, Shuyue; Bit, Subhrangshu; Jasodanand, Varuna H; Liu, Yi; Kolachalama, Vijaya B · Int J Med Inform · 2026

basic_science · Level V

Where this comes from

Abstract

To evaluate whether a tool-using agent-based system built on large language models (LLMs) outperforms standalone LLMs on medical question-answering tasks. We developed a unified, open-source LLM-based agentic system that integrates document retrieval, reranking, evidence grounding, and diagnosis generation to support dynamic, multi-step medical reasoning. Our system features a lightweight retrieval-augmented generation pipeline for efficient evidence retrieval and reranking, coupled with a cache-and-prune memory bank, enabling efficient long-context inference beyond standard LLM limits. The system autonomously invokes specialized tools, eliminating the need for manual prompt engineering or brittle multi-stage templates. We compared the agentic system against standalone language models on various medical question-answering benchmarks. Evaluated on five well-known benchmarks, our system outperforms or closely matches state-of-the-art proprietary and open-source models in multiple-choice and open-ended formats. Specifically, it achieved accuracies of 82.98% on the United States Medical Licensing Examination (USMLE) Step 1 and 86.24% on Step 2, surpassing GPT-4's 80.67% and 81.67%, respectively, while closely matching on Step 3 (88.52% vs. 89.78%). Our findings highlight the value of combining tool-augmented and evidence-grounded reasoning strategies to build reliable and scalable medical artificial intelligence systems.

Medical subject headings