A unified foundation model for heterogeneous EEG signal modeling via language-aligned semi-supervised learning.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42673935.
- Also identified by DOI 10.1016/j.neunet.2026.109544.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Electroencephalography (EEG) signals are highly heterogeneous across datasets, posing challenges for scalable and generalizable representation learning. Existing EEG foundation models predominantly adopt a two-stage paradigm consisting of unsupervised pre-training followed by supervised fine-tuning. This rigid separation struggles to scale to heterogeneous datasets with varying label availability and tends to over-adapt to individual datasets. Moreover, current approaches rarely leverage the semantic structure in label descriptions, limiting their ability to unify supervision across tasks. To address these challenges, we propose USEA, a unified semi-supervised EEG-language alignment approach for heterogeneous EEG signal modeling that formulates EEG modeling as a semantically guided autoregressive representation learning framework. USEA trains a single backbone within a unified optimization framework, enabling joint learning from fully labeled, few-labeled, and unlabeled datasets without relying on task-specific classification heads. When supervision is available, label descriptions are encoded by a frozen language model and used as semantic targets to align EEG representations via cosine similarity. For unlabeled data, the model is trained in a self-supervised manner by autoregressively reconstructing EEG token representations, enabling consistent learning across datasets. To enhance optimization stability and cross-dataset robustness, we adopt a normalization-free Transformer backbone based on Dynamic Tanh (DyT), which mitigates sensitivity to dataset-specific activation statistics while preserving amplitude information critical for EEG signals. We evaluate USEA on nine public EEG datasets spanning diverse paradigms and recording configurations. Experimental results demonstrate competitive performance against state-of-the-art baselines, highlighting the effectiveness of the proposed unified semi-supervised framework for scalable EEG foundation modeling.