Translation readiness of model-based synthetic tabular data in healthcare: a systematic review and governance audit.

Castagno, Simone; Subramanian, Alagu; Epanomeritakis, Ilias E; Gompels, Benjamin; McDonnell, Stephen; Birch, Mark; van der Schaar, Mihaela; McCaskie, Andrew · J Am Med Inform Assoc · 2026

systematic_review · Level I

Where this comes from

Abstract

To evaluate the clinical applications and translation readiness of model-based synthetic tabular data in healthcare, and identify gaps in governance reporting that may hinder translation. We systematically searched Ovid MEDLINE and Embase (2010-August 2025; PROSPERO: CRD42025635514) for studies that generated and applied model-based synthetic tabular data in clinical contexts. Screening used a "human-in-the-loop" large language model workflow alongside independent manual review, achieving 100% sensitivity for included studies. Unlike prior reviews focused primarily on evaluation methodology, we mapped use-cases and deployment paradigms, and audited translation-readiness reporting using a predefined governance framework (validation depth, privacy, fairness, regulatory alignment). Thirty-seven studies (2019-2025) were included. GANs predominated; other approaches included VAEs, diffusion models, LLM-based synthesis, and Bayesian networks. Dataset augmentation was the primary application, often improving downstream model performance for rare outcomes. Emerging applications included synthetic control cohorts and algorithmic bias mitigation. Translation-readiness reporting was limited: 34/37 studies (92%) relied solely on internal validation, 9/37 (24%) used formal privacy models, 6/37 (16%) reported explicit fairness evaluations, and 6/37 (16%) addressed regulatory alignment. Few studies distinguished "no-release" from "delayed-release" paradigms. A systemic gap exists between methodological innovation and deployment-readiness reporting. Model-based synthetic data show clear value for augmentation and class balancing, but inconsistent reporting of validation, privacy, fairness, and regulatory considerations limits confidence in clinical deployment. We propose TRUST-SD (Transparency and Reporting for Utility, Safety, and Translation of Synthetic Data), an author-derived, preliminary, evidence-informed reporting checklist spanning 7 domains, as a starting point for community refinement and consensus-building.