Automated standardization and harmonization of laboratory units in large-scale clinical data using open-source R functions.

Zayed, Ahmed Medhat; Sarikakis, Ilias; Delvaux, Nicolas · Int J Med Inform · 2026

basic_science · Level V

Where this comes from

Abstract

Clinical laboratory data constitute a major component of electronic medical records, yet their secondary use is hindered by inconsistent unit representations, especially in heterogenous multi-source databases. While UCUM and LOINC standards exist, real-world datasets often contain non-standard unit strings and fragmented mappings, limiting interoperability and introducing analytic bias. This study aimed to develop and evaluate open-source functions for automated standardization of units to UCUM-valid formats and harmonization of laboratory results to reference units within LOINC groups. We developed and evaluated two scalable open-source R functions: one for standardizing heterogeneous unit expressions into valid UCUM codes using preprocessing, deterministic mapping, and rule-based correction; and another for harmonizing UCUM-standardized units and values to designated SI or Conventional reference units. Both functions were applied to 163.9 million LOINC-mapped quantitative results from the Intego-II primary care database. Validation involved UCUM web service checks. Among 163,946,741 records with 2,019 distinct unit strings, 157,874,970 (96.2%) were standardized, reducing unit heterogeneity by 81% (from 2,019 to 381 UCUM-valid units). Before applying the function, only 70% of records included UCUM-valid units. Validation confirmed 99.97% correctness of standardized units. Harmonization converted 83.4% of standardized records to reference units within two minutes, consolidating them into 39-42 distinct units per system. Approximately 33% required numeric conversion. All conversions matched UCUM API outputs with 100% concordance rate. Unharmonized cases (16.6%) were primarily due to dimensional mismatches or incompatible arbitrary units. Harmonization also corrected LOINC property inconsistencies in 2.8 million records and reduced code fragmentation, improving distributional characteristics in up to 23% of LOINC groups. Automated, scalable standardization and harmonization of laboratory units are feasible and important, substantially improving consistency and interoperability of large multi-source datasets. These functions enable reliable data integration for analytics and machine learning, reduce bias from coding and unit variability, and provide a quality-control mechanism for identifying structural inconsistencies. The open-source implementation supports integration into routine ETL pipelines and advances the reuse of real-world laboratory data.

Medical subject headings