Machine learning for site risk prediction in clinical trials: development, external validation, and operational application in site qualification.

Yang, Zhiwen · Int J Med Inform · 2026

retrospective_cohort · Level III

Where this comes from

Abstract

Site selection and qualification represent critical operational challenges in clinical trials, particularly in rare diseases like transthyretin amyloid cardiomyopathy (ATTR-CM)-a progressive cardiac disease caused by extracellular deposition of misfolded transthyretin protein in the myocardium affecting <50,000 US patients. Current Site Qualification Risk Assessment (SQRA) approaches-expert judgment, rule-based scoring, or historical performance-lack systematic validation. From a health informatics perspective, clinical trial data stored in disparate systems (eTMF, CTMS, ClinicalTrials.gov) remain underutilized for proactive risk prediction. We developed and validated an ML-based clinical decision support tool for site risk prediction in ATTR-CM trials with external validation in Duchenne muscular dystrophy (DMD), designed for integration into SQRA workflows to guide Clinical Project Managers (CPMs) in site selection and Site Qualification Visit (SQV) decisions. We analyzed 460 sites from 42 ATTR-CM studies and 761 sites from 89 DMD studies using data from ClinicalTrials.gov. A hybrid risk scoring system (expert consensus + LASSO optimization) captured enrollment, data quality, compliance, and complexity dimensions. Four ML algorithms were trained (n = 368) and validated (n = 92) using 5-fold cross-validation and temporal validation. External validation used DMD sites from different disease context and time period. SHAP analysis with bootstrap confidence intervals (1000 iterations) and permutation tests ensured interpretability. We designed an operational workflow integrating the tool into SQRA processes and projected impact based on literature-reported trial parameters. Support Vector Machine achieved 98.91% accuracy (95% CI: 93.8%-99.9%) in ATTR-CM, correctly classifying 91 of 92 sites. Subgroup analysis showed 100% accuracy across all geographic regions and site types, confirming no algorithmic bias. External validation in DMD yielded 81.21% accuracy (95% CI: 73.3%-88.5%), demonstrating cross-disease generalizability. SHAP identified data quality risk (0.112, p < 0.001), screen failure rate (0.090, p < 0.001), and enrollment risk (0.084, p < 0.001) as key predictors. Feature importance rankings showed high consistency between diseases (Spearman ρ = 0.905, p = 0.002). We developed an operational-ready ML-based tool for site risk prediction, achieving near-perfect accuracy in ATTR-CM with demonstrated cross-disease generalizability in DMD. The tool is designed for integration into SQRA workflows to support CPMs in site selection and SQV decisions. Feature importance consistency (Spearman ρ = 0.905) supports applicability across diverse trial portfolios. However, critical limitations remain: lack of validation against ground truth monitoring outcomes, DMD features estimated using proxies, and projected impact not yet validated through operational deployment. This study demonstrates how health informatics can provide decision support for clinical research operations, with pilot implementation essential to validate projected benefits.

Medical subject headings