Psychometric validation of a community healthcare scorecard used for marginalized populations in Bangladesh.

Jubaer, Md Talha; Alam, Md Muhitul; Rayhan, Md Israt; Sultana, Nayeem; Miah, Abu Said Md Juel; Ahsan, Mohd Rubayat · PLoS One · 2026

Where this comes from

Abstract

Achieving equitable and quality healthcare for marginalized groups remains a critical challenge in Bangladesh. Community perceptions are crucial for assessing healthcare system performance, yet tools like Community Scorecards (CSCs) often lack rigorous psychometric validation, limiting their usefulness for evidence-based policy decisions. This study aims to psychometrically validate a 14-item CSC for measuring perceived healthcare service quality among marginalized populations in Bangladesh, using both Classical Test Theory (CTT) and the Generalized Partial Credit Model (GPCM) from Item Response Theory. Data were collected in 2023 from 311 community scoring units across nine marginalized population groups. Each scoring unit represents a Focus Group Discussion (FGD) in which participants reached a group consensus rating on each item using the five-point Likert scale (1 = Very Bad to 5 = Very Good) provided in the questionnaire. We conducted exploratory analysis, factorability checks including the Kaiser-Meyer-Olkin measure and Bartlett's test of sphericity, exploratory factor analysis (EFA) with oblique (promax) rotation, reliability assessment using both Cronbach's alpha and McDonald's omega, and GPCM analysis to evaluate item and scale performance. The Kaiser-Meyer-Olkin measure confirms excellent sampling adequacy (0.83), and Bartlett's test of sphericity was significant (χ² = 1392.90, df = 91, p < 0.001), supporting factorability. EFA reveals a three-factor structure (Accessibility and Fairness, Institutional Responsiveness, Maternal and Child Health Support) explaining 52.7% of variance. The scale shows high internal consistency (Cronbach's α = 0.84; McDonald's ω = 0.87). Although three correlated factors emerge at the CTT level, the GPCM is estimated under a unidimensional assumption empirically supported by a dominant first factor (eigenvalue 4.64 versus 1.58 for the second) and by a moderate average inter-item correlation (0.27). On model fit, the GPCM was preferred over a more restrictive partial credit model on both AIC and BIC, and no item showed systematic residual misfit. GPCM results indicate that all items have significant, positive discrimination parameters (range 0.47 to 2.50), with fairness and access items being the most discriminating. Although the response options were originally designed as a five-point scale, category 5 (Very Good) was not endorsed for any health item, so the GPCM was effectively estimated on the four empirically used categories. Category thresholds were generally well ordered and most transitions were statistically significant, though a small number of intermediate thresholds were not, suggesting that some adjacent categories may be less clearly separated in respondents' minds and warrant attention in future revisions. The Test Information Function shows peak measurement precision for respondents with low to moderate perceived quality levels. The CSC appears to be a psychometrically sound and contextually relevant instrument for the populations studied. Its precision in measuring lower quality perceptions makes it particularly useful for monitoring disparities and evaluating interventions targeting underserved populations. Confirmatory testing in independent samples and additional checks of local independence will further strengthen confidence in its dimensional structure.

Medical subject headings