Defining evidential support for clinical effects in occupational therapy through Bayesian calibration of the paired samples <i>t</i>-test.

Martin, Daniel W · Disabil Rehabil · 2026

other · Level V

Where this comes from

Abstract

Heuristic effect size benchmarks for Cohen's <i>d</i>, Hedges' <i>g</i>, and <i>p</i>-values lack statistical and clinical grounding, producing misleading interpretations of within-subject occupational therapy (OT) research outcomes. This study aimed to recalibrate <i>d</i>, <i>g</i>, and <i>p</i>-value thresholds for paired samples <i>t</i>-tests using Bayesian posterior probabilities representing evidential support for a clinical effect. Using JASP's Bayesian Paired Samples <i>t</i>-test module, Bayes Factors (BF<sub>10</sub>) were iteratively derived across 67 sample sizes (<i>N</i> = 5-500) and converted to <i>d</i>/<i>g</i> thresholds corresponding to posterior probabilities of 50%, 75%, 90%, 95%, and 99%. Sensitivity analyses evaluated threshold stability across Cauchy prior scales. Two generalized additive models compared BF<sub>10</sub> and <i>p</i>-values for predicting effect sizes using 100 published OT studies from the <i>American Journal of Occupational Therapy</i>. Recalibrated thresholds corrected small- and large-sample bias. BF<sub>10</sub> predicted effect sizes more accurately than <i>p</i>-values (<i>R</i><sup>2</sup> = 0.992 vs. 0.966; ΔAIC = 154.81). Thresholds were robust to prior specification. <i>p</i>-values near .05 consistently failed to achieve posterior probabilities of 75% or greater. These recalibrated thresholds provide occupational therapists a principled, sample size-sensitive tool for interpreting within-subject outcomes probabilistically, offering a practical alternative to conventional heuristic benchmarks.