IDQuAD: Infectious disease question and answering dataset.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41066458.
- Also identified by DOI 10.1371/journal.pone.0333075 and PMC identifier 12510504.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
While large language models (LLMs) have made significant advances in various fields, The study of applying LLMs to infectious disease-specific tasks has lagged behind. This study addresses this gap by introducing the Infectious Disease Question and Answering Dataset (IDQuAD), which is a novel dataset designed to train and evaluate LLMs in infectious disease-related queries. IDQuAD is constructed using medical papers, patents, and news, and employs innovative methodologies such as generating answers before questions and using counterfactual thinking to enhance the quality of the Question Answering (QA) pairs. In the experimental phase, we fine-tuned the Mistral-7B model on the IDQuAD dataset to test the effectiveness of our proposed datasets on LLM performance in QA tasks related to infectious diseases. The fine-tuned Mistral-7B model demonstrated substantial performance improvements, with its EM score increasing from 28.49% to 65.47% in the one-shot setting. Additionally, we evaluated other LLMs across various setups. Among all models tested, our fine-tuned model achieved the highest performance across metrics and settings. In conclusion, this study introduces IDQuAD as a foundational dataset for infectious disease research, demonstrating the effectiveness of fine-tuning LLMs and paving the way for future advances in dataset development and LLM refinement for infectious disease tasks.
Medical subject headings
- Communicable Diseases
- Language