Impact of quality scales on levels of evidence inferred from a systematic review of exercise therapy and low back pain.

Colle, Florence; Rannou, François; Revel, Michel; Fermanian, Jacques; Poiraudeau, Serge · Arch Phys Med Rehabil · 2002

systematic_review · Level I

Where this comes from

Abstract

To assess whether the scale used affects levels of evidence inferred from a systematic review of studies on exercise therapy and chronic low back pain (LBP). Twenty trials previously analyzed in a systematic review were assessed by 2 readers using 16 different scales. Tertiary care teaching hospital in France. Chronic LBP patients. Not applicable. For the scales allowing classification into high- and low-quality trials, a rating system with 4 levels of evidence was used to summarize conclusions drawn. The Spearman rank correlation coefficient was used to assess correlations between the scores obtained with the different scales. Interrater reliability of the scales was assessed with the intraclass correlation coefficient and the Bland and Altman method, and the degree of agreement between the readers was calculated using the kappa coefficient. Two of the 3 main results of the systematic review (conflicting evidence on the effectiveness of exercise therapy compared with inactive treatments; strong evidence that exercise therapy is more effective than usual care by a general practitioner) were influenced by the scale used. The range of the Spearman rank correlation coefficients between the different scales was wide (range,.49-.94), the interreader reliability of the scales was heterogeneous, and the interreader agreement was often low (kappa<or=.60 for 7/10 tests). The use of summary scores to identify physical therapy trials of high quality is questionable. Different quality assessment scales should probably be used to assess pharmacologic interventions and physical therapies. Development and validation of quality scales specific to physical treatments are needed.

Medical subject headings