Detecting Uncoded Self-Harm in Veterans' Electronic Health Records Using Positive and Unlabeled Learning: Retrospective Observational Study.
retrospective_cohort · Level III
Where this comes from
- Record sourced from PubMed, PMID 42001435.
- Also identified by DOI 10.2196/89071.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Suicide and self-harm remain major public health concerns in the United States. Early identification is critical for effective intervention, yet underdiagnosis and undercoding are common across mental health conditions, and only positive cases are typically labeled in healthcare data. As a result, reliable negative examples are missing. Positive and Unlabeled (PU) learning is well-suited to such data, enabling estimation of phenotype prevalence and identification of undiagnosed individuals at elevated risk for self-harm as well as other mental illnesses. To identify U.S. Veterans whose self-harm events were not explicitly captured through diagnostic codes in electronic health records (EHRs) and estimate the prevalence of ever self-harm cases among Veterans using a novel PU learning algorithm applicable to undetected mental health diagnoses. We performed a retrospective observational study using Veterans Health Administration EHRs (from October 1, 1999 to August 31, 2019), selecting a random 25% sample of 1,329,120 Veterans out of 5,316,480 (1,193,563 males and 135,557 females) with at least 2 years of observation. The study cohort comprised 24,625 veterans with coded self-harm and 1,304,495 uncoded for self-harm, with the mean age for coded individuals 38.39 (SD 12.17) and uncoded individuals 48.76 (SD 15.04). We applied our PULSNAR (Positive Unlabeled Learning Selected Not At Random) algorithm to estimate the proportion of individuals with uncoded self-harm. The selected covariates included age at enrollment and the presence or absence of recorded medical conditions, procedures, and clinical observations throughout the observation period. Four experts (raters) independently reviewed charts of 97 uncoded Veterans, each selected from 1% intervals of calibrated PULSNAR probabilities from 0.01 to 0.97. Agreement was assessed among raters, PULSNAR classifications, and consensus review decisions. Post hoc calibration was used to refine prevalence estimates. Of the 159,049 covariates in the dataset, the XGBoost model within our PULSNAR framework identified 1,302 (0.82%) as informative for classification. Only 1.85% (24,625/1,329,120) of Veterans had diagnostic codes indicating self-harm events, while PULSNAR estimated an overall prevalence of 10.46% (139,026/1,329,120) by identifying an additional ?=8.77% (114,404/1,304,495) of self-harm cases among the uncoded population. Of the 97 chart-reviewed patients, 39 had documented but uncoded self-harm. PULSNAR estimates were post hoc calibrated such that their sum over the 97 cases equaled 39, which resulted in PULSNAR adjusted coded and imputed estimation of 7.91% (105,133/1,329,120). When applied to the 1.3M Veterans, PULSNAR suggests that coded self-harm represents only 23.4% (95% CI: 17.76% to 31.51%) of all documented (coded + notes) self-harm. Under the Selected Not At Random assumption, PULSNAR provides an innovative and scalable framework for estimating the clinically documented prevalence of mental health conditions and identifying the uncoded individuals with calibrated prediction, without requiring confirmed negative labels. This method offers an alternative to time-consuming chart reviews for detecting likely cases missing structured coding capture. By addressing diagnostic undercoding of mental health conditions in EHRs, this approach has the potential to enhance the estimation of mental health prevalence and support screening, activation of automated clinical decision support, targeted intervention, better resource allocation, and research to improve outcomes in real-world settings.