Hidden Markov model using Dirichlet process for de-identification.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 26407642.
- Also identified by DOI 10.1016/j.jbi.2015.09.004 and PMC identifier 4984397.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
For the 2014 i2b2/UTHealth de-identification challenge, we introduced a new non-parametric Bayesian hidden Markov model using a Dirichlet process (HMM-DP). The model intends to reduce task-specific feature engineering and to generalize well to new data. In the challenge we developed a variational method to learn the model and an efficient approximation algorithm for prediction. To accommodate out-of-vocabulary words, we designed a number of feature functions to model such words. The results show the model is capable of understanding local context cues to make correct predictions without manual feature engineering and performs as accurately as state-of-the-art conditional random field models in a number of categories. To incorporate long-range and cross-document context cues, we developed a skip-chain conditional random field model to align the results produced by HMM-DP, which further improved the performance.
Medical subject headings
- Computer Security
- Confidentiality
- Electronic Health Records
- Narration
- Natural Language Processing
- Pattern Recognition, Automated