Identifying family structures from obituaries and matching them to patients in an electronic heath record.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40576211.
- Also identified by DOI 10.1093/jamia/ocaf102 and PMC identifier 12361849.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Family data are a valuable data source in bioinformatic research. This is because family members often share common genetic and environmental exposures. Collecting this family data is traditionally very labor intensive but advances in electronic health record (EHR) data mining has proven useful when identifying pedigrees linked to longitudinal health histories. These are called e-pedigrees. Unfortunately, e-pedigrees tend to miss the oldest patients who inherently have the longest and richest health histories. A good source of family data from older generations includes obituaries, as they have a formulaic nature making them a good candidate for natural language processing (NLP) that can extract relationships to the decedent. While there have been several studies on obtaining such data from obituaries, we demonstrate for the first time approaches that tie that information to an EHR. Natural language processing extraction resulted in 8 166 534 family members being abstracted from 567 279 obituaries published in the state of Wisconsin. After matching decedent and family members to patients in the EHR, we identified 200 033 unique patients that were put in 53 640 pedigrees. The largest pedigree consisted of 21 individuals. Heritability of adult height was quantified (H2=0.51±0.04, P<1.00e-07) demonstrating these data's use in genetic research. The heritability data, coupled with overlapping data in a biobank, suggested 80%-90% of familial relationships were accurately defined. The totality of these findings demonstrate obituaries with the oldest people in society can be highly informative for bioinformatic research. Code is available on GitHub at https://github.com/jgmayer672/ObituaryNLP.
Medical subject headings
- Electronic Health Records
- Natural Language Processing
- Data Mining
- Pedigree
- Family