Evaluating the clinical readiness of artificial intelligence in EEG-based epilepsy diagnosis.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 41330046.
- Also identified by DOI 10.1088/1741-2552/ae2714.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
<i>Objective.</i>Automated electroencephalography (EEG)-based epilepsy diagnosis has reported near-perfect accuracies for almost two decades on a benchmark dataset, yet virtually no system is used in routine care. We critically re-examined this translation gap by reproducing five widely cited Artificial Intelligence (AI) models spanning statistical feature extraction, classical machine-learning and deep-learning paradigms, and assessed their ability to generalise from the benchmark dataset to a newly curated, clinically verified scalp-EEG cohort.<i>Approach.</i>All models were implemented as originally described and trained with ten-fold cross-validation on the benchmark dataset. External validation was performed on our independently curated dataset made publicly available, comprising 30 subjects (15 epilepsy, 15 healthy) recorded with 19-channel scalp EEG under standard clinical protocols. We further examined the influence of subject-level data leakage by contrasting performance when training/testing samples overlapped with those when completely independent patient partitions were enforced. Accuracy, sensitivity, specificity, and AUC of ROC (with 95% Confidence Intervals) were the primary metrics.<i>Main results.</i>When transferred unchanged to the external cohort, overall accuracy fell from⩾94% on the benchmark dataset to 42%-53%, with sensitivities as low as 0.97% for the deep convolutional neural network and specificities dropping to 4% for time-frequency methods. Permitting subject overlap artificially elevated accuracy to 59%-96%, whereas strict patient separation reduced it to 41%-53% in 95% Confidence Interval. Deep-learning models exhibited the steepest decline, confirming over-fitting to subject-specific artefacts. Statistical feature-based approaches, though less affected, still under-performed clinically acceptable thresholds.<i>Significance.</i>Our results expose key translational barriers in AI for EEG-based epilepsy diagnosis-data leakage, acquisition bias, and overfitting to patient idiosyncrasies-leading to severe performance erosion on clinical data. Rigorous patient-independent validation, transparent reporting (aligned with CARE principles), and well-curated multi-channel scalp-EEG datasets are essential to ensure clinically dependable AI tools for epilepsy diagnosis.
Medical subject headings
- Electroencephalography
- Epilepsy
- Artificial Intelligence