Agile text mining for the 2014 i2b2/UTHealth Cardiac risk factors challenge.
other
Where this comes from
- Record sourced from PubMed, PMID 26209007.
- Also identified by DOI 10.1016/j.jbi.2015.06.030 and PMC identifier 4737484.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
This paper describes the use of an agile text mining platform (Linguamatics' Interactive Information Extraction Platform, I2E) to extract document-level cardiac risk factors in patient records as defined in the i2b2/UTHealth 2014 challenge. The approach uses a data-driven rule-based methodology with the addition of a simple supervised classifier. We demonstrate that agile text mining allows for rapid optimization of extraction strategies, while post-processing can leverage annotation guidelines, corpus statistics and logic inferred from the gold standard data. We also show how data imbalance in a training set affects performance. Evaluation of this approach on the test data gave an F-Score of 91.7%, one percent behind the top performing system.
Medical subject headings
- Cardiovascular Diseases
- Data Mining
- Diabetes Complications
- Electronic Health Records
- Narration
- Natural Language Processing