Data Augmentation for Few-Shot Biomedical NER Using ChatGPT.

Mu, Wenxuan; Zhao, Di; Meng, Jiana; Chen, Peng; Sun, Shichang; Yang, Yumeng; Wang, Jian; Lin, Hongfei · Artif Intell Med · 2026

basic_science · Level V

Where this comes from

Abstract

Data Augmentation (DA) aims to create a new dataset to address the lack of data in various domains. Particularly in few-shot scenarios of the biomedical Named Entity Recognition (NER) domain, an effective DA method can enhance data diversity, reduce overfitting, and significantly improve the model's generalization ability. In this work, we propose a novel DA method for NER tasks, which uses ChatGPT and prompt learning to extract high-quality data from large language models. The entity recognition tasks are then performed via transfer learning and efficient decoding strategies. Moreover, this study conducted extensive experiments on four publicly available biomedical datasets (BC5CDR, NCBI, BioNLP11EPI, and BioNLP13GE), demonstrating that our methods exhibit strong stability and entity recognition capabilities even in extremely limited scenarios. In the 5-shot, 20-shot, and 50-shot scenarios, the average F1 scores of the four datasets reached 72.96%, 75.05%, and 77.42%, respectively.

Medical subject headings