Adapting existing natural language processing resources for cardiovascular risk factors identification in clinical notes.
Where this comes from
- Record sourced from PubMed, PMID 26318122.
- Also identified by DOI 10.1016/j.jbi.2015.08.002 and PMC identifier 4983192.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
The 2014 i2b2 natural language processing shared task focused on identifying cardiovascular risk factors such as high blood pressure, high cholesterol levels, obesity and smoking status among other factors found in health records of diabetic patients. In addition, the task involved detecting medications, and time information associated with the extracted data. This paper presents the development and evaluation of a natural language processing (NLP) application conceived for this i2b2 shared task. For increased efficiency, the application main components were adapted from two existing NLP tools implemented in the Apache UIMA framework: Textractor (for dictionary-based lookup) and cTAKES (for preprocessing and smoking status detection). The application achieved a final (micro-averaged) F1-measure of 87.5% on the final evaluation test set. Our attempt was mostly based on existing tools adapted with minimal changes and allowed for satisfying performance with limited development efforts.
Medical subject headings
- Cardiovascular Diseases
- Data Mining
- Diabetes Complications
- Electronic Health Records
- Narration
- Natural Language Processing