Development and validation of data-driven, decision tree-based algorithms for identifying Behçet's disease in claims data.

Sada, Ken-Ei; Miyawaki, Yoshia; Yanai, Ryo; Kida, Takashi; Onishi, Akira; Yoshimi, Ryusuke; Ichinose, Kunihiro; Shimojima, Yasuhiro · Int J Med Inform · 2026

cross_sectional · Level IV

Where this comes from

Abstract

To develop and externally validate novel, data-driven algorithms that are based on appropriate variable selection methods for identifying patients with Behçet's disease in Japan. This retrospective cross-sectional study included 13,538 patients from six tertiary hospitals (November-December 2023). One year of claims data was linked to chart-confirmed Behçet's disease diagnoses. Patients were randomly divided into training (n = 8,811) and test (n = 3,775) sets, with external validation (n = 952) from another hospital. Feature selection among Behçet's disease-coded patients used the Least Absolute Shrinkage and Selection Operator, Boruta, and Recursive Feature Elimination. The diagnostic performance of the rule-based algorithms, which were derived from the decision tree models, was evaluated using accuracy, sensitivity, specificity, positive predictive value (PPV), negative predictive value, and F1 score. Diagnosis codes alone achieved high sensitivity (1.000) and specificity (0.992) but modest PPV (0.767, test set; 0.850, external validation). Incorporating sulphamethoxazole-trimethoprim and colchicine prescriptions improved the positive predictive value, which was 0.793 in the test set and 0.865 in external validation. Incorporating prescriptions alongside diagnosis codes improved PPV while maintaining high sensitivity and specificity. Building upon a data-driven framework that integrates variable selection methods and decision tree analysis, this study provides a validated and scalable approach for reliable claims-based research on Behçet's disease.

Medical subject headings