Learning from experts: A self-improving LLM framework for study population generation in clinical research.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 41197329.
- Also identified by DOI 10.1016/j.ijmedinf.2025.106171.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
The widespread adoption of electronic health records has led to the rapid accumulation of real-world data (RWD), an essential basis for generating real-world evidence (RWE). While large language models (LLMs) have supported multiple stages of RWD-driven research, their application to the design of study populations with both interpretability and credibility remains a challenge, which serves as a bridging role between study objectives and downstream analyses. In this study, we propose CriteriaLLM, a framework that enables LLMs to generate eligible study populations directly from clinical research objectives by incorporating clinician feedback. Inspired by the after-action review method, which facilitates learning from past experiences and feedback, we build an expert knowledge base that records the LLM output study populations and the modifications made by the clinician. A dual-retrieval algorithm, combining disease domain relevance and lexical similarity, then identifies relevant historical cases from the expert knowledge base to guide future generations. To ensure clinical relevance and real-world applicability, we introduce a continuous validation loop where expert feedback is iteratively integrated, refining model performance over time. We evaluated our framework on 254 published clinical studies based on the MIMIC-III database using four representative LLMs: GPT-4o, Deepseek-R1, and two LLaMA models. The experimental results indicate that the proposed framework could effectively generate a high-quality study population with the highest Macro F1 score on 0.9180, maintaining generalizability across foundation models with varying parameter sizes and deployment methods. Our expert-in-the-loop framework allows LLMs to generate eligible study populations from clinical objectives without additional fine-tuning. By integrating structured expert feedback and retrieval guidance, it enhances the quality and reliability of study criteria. With continuous validation, the framework highlights a scalable approach toward self-improving systems that bridge generative AI with the demands for clinical appropriateness, reliability, and interpretability in clinical research.
Medical subject headings
- Electronic Health Records
- Biomedical Research
- Natural Language Processing
- Language