Quality assessment of YouTube videos on total knee arthroplasty: a comparative evaluation of orthopedic surgeons and artificial intelligence models.
cross_sectional · Level IV
Where this comes from
- Record sourced from PubMed, PMID 42296690.
- Also identified by DOI 10.1016/j.knee.2026.104529.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
YouTube is widely used by patients seeking preoperative information before total knee arthroplasty (TKA). As online video content may influence patient expectations and postoperative satisfaction, reliable tools are needed to assess educational quality. Advances in artificial intelligence (AI) enable large-scale evaluation of digital medical content. This study assessed the educational quality of the most viewed YouTube videos on TKA and compared evaluations by orthopedic surgeons and AI models. On December 10, 2025, YouTube was searched using the term "total knee arthroplasty." The 61 most viewed videos were screened; 50 met inclusion criteria. Two orthopedic surgeons independently evaluated each video using DISCERN, JAMA benchmark criteria, Global Quality Score (GQS), Patient Education Materials Assessment Tool (PEMAT), and a novel Total Knee Arthroplasty-Specific Scoring System (TKA-SS). Interobserver reliability was calculated using intraclass correlation coefficients (ICC). AI evaluations were performed using transcript-based and multimodal large language models. Correlations between human and AI-generated scores and absolute differences were analyzed. Educational quality ranged from low to moderate with substantial heterogeneity. Interobserver reliability was excellent for most instruments, particularly DISCERN (ICC = 0.985), but low for PEMAT-understandability. Significant positive correlations were observed between surgeon reference and AI-generated scores across all instruments (all p < 0.001), strongest for TKA-SS. Mean absolute differences were small, and no significant differences were found between AI models. Widely viewed TKA-related YouTube videos demonstrate predominantly low-to-moderate educational quality. AI-generated assessments show meaningful agreement with surgeon evaluations and may serve as a scalable tool for preliminary content stratification rather than definitive quality assessment.