Assessing the fidelity of synthetic relational databases with high-dimensional categorical data: Proposal of a scalable visual framework.
other
Where this comes from
- Record sourced from PubMed, PMID 42700765.
- Also identified by DOI 10.1016/j.ijmedinf.2026.106711.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Although data reuse is increasingly common in healthcare research, changing regulatory frameworks are impeding these efforts. Synthetic data constitute a promising issue but also create challenges, such as the assessment of data fidelity and the ability to output identical statistical results. Conventional validation approaches often rely on univariate or bivariate comparisons, which fail to capture the complexity of associations between multivalued categorical variables. We built and studied a fictitious database of 10,000 hospital stays reproducing the structure of the French Programme de Médicalisation des Systèmes d'Information database. Each stay included single-valued variables (one value per individual: sex, age in deciles, and diagnosis-related group) and multivalued variables (zero, one or several values per individual: diagnoses coded according to the International Classification of Diseases, 10th Edition, and procedures coded according to the French Classification Commune des Actes Médicaux). All categorical variables were binarized, and thousands of pairwise association metrics (primarily odds ratios) were calculated for the reference and evaluation datasets. The results were summarized using curves, bubble charts, heatmaps, and coefficients such as exponential mean deviation. Simulated data degradations from 0% to 100% were introduced to evaluate the method's sensitivity. We analyzed 500 ICD-10 diagnoses and 450 CCAM procedures, representing 225,000 possible combinations. We developed and evaluated graphical representations for assessing data fidelity at a glance. In simulations of an increasing degree of data degradation, those graphical representations and comprehensive, quantitative metrics facilitated the detection of the gradual loss of data fidelity. We developed a simple, scalable, agnostic framework for assessing the fidelity of healthcare databases by systematically analyzing associations among the modalities of coded variables. This method complements existing approaches. It is particularly suitable for the evaluation of synthetic relational databases because it offers both general and granular insights into data fidelity loss.