Machine Listening for OSA Diagnosis: A Bayesian Meta-Analysis.
meta_analysis · Level I
Where this comes from
- Record sourced from PubMed, PMID 40220991.
- Also identified by DOI 10.1016/j.chest.2025.04.006 and PMC identifier 12405919.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Among 1 billion patients worldwide with OSA, 90% remain undiagnosed. The main barrier to diagnosis is the overnight polysomnogram, which requires specialized equipment, skilled technicians, and inpatient beds available only in tertiary sleep centers. Recent advances in artificial intelligence (AI) have enabled OSA detection using breathing sound recordings. What is the diagnostic accuracy of and how can we optimize machine listening for OSA? PubMed, Embase, Scopus, Web of Science, and IEEE Xplore databases were systematically searched. Two masked reviewers selected studies comparing the patient-level diagnostic performance of AI approaches using overnight audio recordings vs conventional diagnosis (apnea-hypopnea index) using a train-test split or k-fold cross-validation. Bayesian bivariate meta-analysis and meta-regression were performed. Publication bias was assessed by using a selection model. Risk of bias and evidence quality were assessed by using the Quality Assessment of Diagnostic Accuracy Studies-2 and the Grading of Recommendations, Assessment, Development, and Evaluation tools. From 6,254 records, 16 studies (41 models) trained on 4,864 participants and tested on 2,370 participants were included. No study had a high risk of bias. Machine listening achieved a pooled sensitivity (95% credible interval) of 90.3% (86.9%-93.1%), a specificity of 86.7% (83.1%-89.7%), a diagnostic OR of 60.8 (39.4-99.9), and positive and negative likelihood ratios of 6.78 (5.34-8.85) and 0.113 (0.079-0.152), respectively. At apnea-hypopnea index cutoffs of ≥ 5, ≥ 15, and ≥ 30 events per hour, sensitivities were 94.3% (90.3%-96.8%), 86.3% (80.1%-90.9%), and 86.3% (79.2%-91.1%); and specificities were 78.5% (68.0%-86.9%), 87.3% (81.8%-91.3%), and 89.5% (84.8%-93.3%). Meta-regression identified increased sensitivity for the following: higher audio sampling frequencies, non-contact microphones, higher OSA prevalence, and train-test split model evaluation. Accuracy was equal regardless of home smartphone vs in-laboratory professional microphone recordings, deep learning vs traditional machine learning, and variations in age and sex. Publication bias was not evident, and the evidence was of high quality. In this study, machine listening achieved excellent diagnostic accuracy, superior to the STOP-Bang (snoring, tiredness, observed apnea, BP, BMI, age, neck size, gender) questionnaire and comparable to common home sleep tests. Digital medicine should be further explored and externally validated for accessible and equitable OSA diagnosis. PROSPERO database; No.: CRD42024534235; URL: https://www.crd.york.ac.uk/PROSPERO/).
Medical subject headings
- Sleep Apnea, Obstructive
- Polysomnography
- Artificial Intelligence