Evaluation of temporal preservation in synthetic longitudinal patient data.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 42372998.
- Also identified by DOI 10.1016/j.jbi.2026.105074.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
This study introduces a set of metrics for evaluating temporal preservation in synthetic longitudinal patient data, defined as artificially generated data that mimic real patients' repeated measurements over time. The proposed metrics assess how synthetic data reproduce key temporal characteristics, categorized into marginal, covariance, individual-level and measurement structures. Strong marginal-level resemblance may be observed even when the covariance structures and individual trajectories are substantially different. Temporal preservation is influenced by factors such as original data quality, measurement frequency, and preprocessing strategies, including binning, variable encoding and precision. Variables with sparse or highly irregular measurement times provide limited information for learning temporal dependencies, yielding reduced resemblance between the synthetic and original data. No single metric adequately captures temporal preservation; instead, a multidimensional evaluation across all characteristics provides a more comprehensive assessment of synthetic data quality. Overall, the proposed metrics elucidate how and why temporal structures are preserved or degraded, enabling more reliable evaluation and improvement of generative models and supporting the creation of temporally realistic synthetic longitudinal patient data.