SNOMED CT entity linking challenge.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 40657868.
- Also identified by DOI 10.1093/jamia/ocaf104 and PMC identifier 12361850.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
This paper presents the results from a competition challenging participants to develop entity linking models using a subset of annotated MIMIC-IV-Note data and the SNOMED CT Terminology. As a basis for this work, a large set of 74 808 annotations was curated across 272 discharge notes spanning 6624 unique clinical concepts. Submissions were evaluated using the mean Intersection-over-Union metric, evaluated at the character level with the 3 best performing solutions awarded a cash prize. The winning solutions employed contrasting approaches: a dictionary-based method, an encoder-based method, and a decoder-based method. Our analysis reveals that concept frequency in training data significantly impacts model performance, with rare concepts proving particularly challenging. High concept entropy and annotation ambiguity were also associated with decreased performance. Findings from this work suggest that future projects should focus on improving entity linking for rare concepts and developing methods to better leverage contextual information when training examples are scarce.
Medical subject headings
- Systematized Nomenclature of Medicine
- Electronic Health Records