Measurement Properties of Delirium Tools in Pediatric Intensive Care: A Systematic Review.
systematic_review · Level I
Where this comes from
- Record sourced from PubMed, PMID 42680167.
- Also identified by DOI 10.1542/peds.2026-077184.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Delirium is common in pediatric intensive care units. Reliable tools are needed, but evidence on measurement properties remains fragmented. To identify pediatric delirium tools and evaluate their measurement properties and certainty of evidence. MEDLINE, EMBASE, PsycINFO, CINAHL, Cochrane Library, and Web of Science were searched without language or date restrictions. Original observational, cross-sectional, and validation studies evaluating at least 1 measurement property of a pediatric delirium tool in acute care were eligible. Two reviewers independently screened studies, extracted data, and assessed risk of bias using the adapted COnsensus-based Standards for the selection of health Measurement INstruments (COSMIN) methods and Quality Assessment of Diagnostic Accuracy Studies-2 for diagnostic accuracy. Evidence was rated using COSMIN-adapted Grading of Recommendations Assessment, Development and Evaluation. From 8378 records, 41 studies were included, covering 15 language versions of 6 tools: Cornell Assessment of Pediatric Delirium (CAPD), preschool-Confusion Assessment Method for Intensive Care Unit, pediatric-Confusion Assessment Method for Intensive Care Unit, Sophia Observation withdrawal Symptoms scale - Pediatric Delirium (SOS-PD), PEdiatric Delirium Scale (PEDS), and Child Delirium Assessment Scale (CDAS). CAPD had the largest evidence base. High-certainty evidence supported criterion validity of CAPD and Confusion Assessment Method tools, with moderate-certainty evidence for reliability. SOS-PD showed moderate-certainty evidence for criterion validity and high-certainty evidence for convergent validity and reliability but was evaluated in fewer studies. Evidence for CDAS and PEDS was limited to single studies. Evidence was uneven across tools and mainly addressed criterion validity and interrater reliability. Measurement error, cross-cultural validity, and subgroup performance were rarely assessed. Criterion validity is difficult because no perfect gold standard exists for diagnosis. SOS-PD showed the most favorable certainty profile across evaluated measurement properties but was evaluated in fewer studies than CAPD. Further studies should address measurement error, cross-cultural validity, subgroup performance, and implementation.