A comprehensive survey on medical concept normalization: Datasets, techniques, applications, and future directions.
review · Level V
Where this comes from
- Record sourced from PubMed, PMID 41780734.
- Also identified by DOI 10.1016/j.jbi.2026.105005.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Medical concept normalization (MCN) involves mapping informal medical terms or phrases to standardized medical concepts, serving a crucial role in medical text analysis, healthcare intelligence, and other applications. Previous studies have primarily focused on developing datasets, refining machine learning algorithms, and applying them to downstream tasks, but a comprehensive survey covering all key aspects of MCN has been lacking. This survey addresses this gap by providing a complete overview of MCN, including task definitions, annotation schemes, datasets, normalization techniques, state-of-the-art performance, and applications. A comparative analysis of benchmark datasets and leading normalization methods is conducted, along with a quantitative evaluation of CNN-based, RNN-based, transformer-based, GNN-based, LLM-based, and other models on popular benchmarks. Guidelines are also provided for selecting the most effective models and techniques. In addition, the applications of MCN in electronic health records (EHRs) management, clinical decision support (CDS), precision medicine, health information exchange (HIE), and others are explored. Finally, key challenges and potential directions for future research are discussed. This survey offers the first comprehensive review of the core components of MCN, serving as a valuable resource for both researchers and practitioners. Resources related to this survey can be accessed on GitHub at: https://github.com/haihua0913/awesome-mcn.