Evaluating the clinical readiness of artificial intelligence in EEG-based epilepsy diagnosis.

Bandhu Lahiri, Jeet; Agarwal, Puneet; Kushwaha, Suman; Singh, Mridula; Panwar, Siddharth · J Neural Eng · 2025

other · Level V

Where this comes from

Abstract

<i>Objective.</i>Automated electroencephalography (EEG)-based epilepsy diagnosis has reported near-perfect accuracies for almost two decades on a benchmark dataset, yet virtually no system is used in routine care. We critically re-examined this translation gap by reproducing five widely cited Artificial Intelligence (AI) models spanning statistical feature extraction, classical machine-learning and deep-learning paradigms, and assessed their ability to generalise from the benchmark dataset to a newly curated, clinically verified scalp-EEG cohort.<i>Approach.</i>All models were implemented as originally described and trained with ten-fold cross-validation on the benchmark dataset. External validation was performed on our independently curated dataset made publicly available, comprising 30 subjects (15 epilepsy, 15 healthy) recorded with 19-channel scalp EEG under standard clinical protocols. We further examined the influence of subject-level data leakage by contrasting performance when training/testing samples overlapped with those when completely independent patient partitions were enforced. Accuracy, sensitivity, specificity, and AUC of ROC (with 95% Confidence Intervals) were the primary metrics.<i>Main results.</i>When transferred unchanged to the external cohort, overall accuracy fell from⩾94% on the benchmark dataset to 42%-53%, with sensitivities as low as 0.97% for the deep convolutional neural network and specificities dropping to 4% for time-frequency methods. Permitting subject overlap artificially elevated accuracy to 59%-96%, whereas strict patient separation reduced it to 41%-53% in 95% Confidence Interval. Deep-learning models exhibited the steepest decline, confirming over-fitting to subject-specific artefacts. Statistical feature-based approaches, though less affected, still under-performed clinically acceptable thresholds.<i>Significance.</i>Our results expose key translational barriers in AI for EEG-based epilepsy diagnosis-data leakage, acquisition bias, and overfitting to patient idiosyncrasies-leading to severe performance erosion on clinical data. Rigorous patient-independent validation, transparent reporting (aligned with CARE principles), and well-curated multi-channel scalp-EEG datasets are essential to ensure clinically dependable AI tools for epilepsy diagnosis.

Medical subject headings