Data Augmentation for Few-Shot Biomedical NER Using ChatGPT.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41338038.
- Also identified by DOI 10.1016/j.artmed.2025.103314.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Data Augmentation (DA) aims to create a new dataset to address the lack of data in various domains. Particularly in few-shot scenarios of the biomedical Named Entity Recognition (NER) domain, an effective DA method can enhance data diversity, reduce overfitting, and significantly improve the model's generalization ability. In this work, we propose a novel DA method for NER tasks, which uses ChatGPT and prompt learning to extract high-quality data from large language models. The entity recognition tasks are then performed via transfer learning and efficient decoding strategies. Moreover, this study conducted extensive experiments on four publicly available biomedical datasets (BC5CDR, NCBI, BioNLP11EPI, and BioNLP13GE), demonstrating that our methods exhibit strong stability and entity recognition capabilities even in extremely limited scenarios. In the 5-shot, 20-shot, and 50-shot scenarios, the average F1 scores of the four datasets reached 72.96%, 75.05%, and 77.42%, respectively.
Medical subject headings
- Data Mining
- Natural Language Processing