ViSTA-SleepNet: A View-Integrated and Subject-Adaptive Transformer Framework for Multimodal Sleep Staging.
Where this comes from
- Record sourced from PubMed, PMID 42678829.
- Also identified by DOI 10.1109/JBHI.2026.3729590.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Accurate sleep staging is essential for the diagnosis and evaluation of sleep disorders. However, convolutional neural network (CNN)-based methods are limited in capturing stage-specific sleep waveforms, and classification errors often occur in transitional sleep regions. In addition, existing models do not adequately account for inter-individual variability related to physiological factors such as age and sex. To address these issues, we propose ViSTA-SleepNet, a two-stage, multi-view, multimodal framework for sleep staging.Specifically, data augmentation is applied to sleep stage transition segments to improve boundary learning. For EEG signals, a multi-view representation learning strategy and a cross-view Transformer are used to jointly model raw waveforms, amplitude variations, and temporally structured events, thereby enhancing the detection of key sleep events. EOG and EMG are modeled as complementary modalities and adaptively fused with EEG features. In addition, a feature-wise linear modulation (FiLM) mechanism based on demographic and physiological information is introduced to support individualized representation learning and joint detection of sleep stages, spindles, and slow waves. Finally, a bidirectional gated recurrent unit (Bi-GRU) and conditional random field (CRF) are combined to improve the physiological consistency of sleep stage transitions.Experiments on Sleep-EDF-20, Sleep-EDF-78, and SHHS using strict 20-fold subject-independent cross-validation show that ViSTA-SleepNet achieves accuracies of 88.5%, 84.5%, and 89.6% for sleep staging, respectively, while exceeding 95% accuracy in spindle and slow-wave detection. Visualization results further confirm its advantages in discriminative performance and physiological plausibility.