EEG-PRIME: Prototype-Aligned Representation Learning with Multi-Level Conditioning for EEG Decoding

arXiv:2608.13072 · cs.AI · Submitted 2026-08-13 · Read on arXiv

Shuailei Zhang, Muyun Jiang, Wei Zhang, Jinbo Chen, Zhiwei Guo, Yong Li, Yi Ding, Cuntai Guan

Nanyang Technological University · Southeast University

cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/ZhangShuailei/EEG-PRIME

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 100/100

The gist: EEG-PRIME is a two-stage EEG foundation model for cross-dataset multi-task EEG decoding, proposed to address the poor generalization of EEG decoding models across datasets and subjects due to domain

Terminology

Summary

EEG-PRIME is a two-stage EEG foundation model for cross-dataset multi-task EEG decoding, proposed to address the poor generalization of EEG decoding models across datasets and subjects due to domain shifts in acquisition protocols and individual neurophysiology. The model combines masked pretraining with prototype-aligned instruction tuning under multi-level conditioning, enabling instruction-aware, subject-invariant EEG decoding across diverse brain-computer interface (BCI) paradigms.

In the pretraining stage, an EEG encoder learns transferable representations via self-supervised masked reconstruction with frequency-cutoff spectral augmentation. The encoder consists of a CNN-based tokenizer that segments EEG into non-overlapping temporal windows and extracts per-window feature vectors via depthwise-separable convolutions, followed by a Transformer that models temporal dependencies across tokens. A random token mask ratio of 0.5 is applied, and the pretraining loss is the mean squared error between reconstructed and original EEG signal values at masked token positions. Pretraining data are unified to a 65-channel 10–10 electrode layout and resampled to 200 Hz.

In the instruction tuning stage, three conditioning levels are introduced: (1) a task-semantic prompt derived from a natural language instruction describing the decoding objective, encoded by a frozen Sentence-BERT (all-mpnet-base-v2) text encoder into a fixed-dimensional embedding; (2) a dataset-level soft embedding that is jointly learned during training and additively combined with the task prompt to capture dataset-specific distributional characteristics; and (3) a subject-invariance constraint enforced via gradient reversal adversarial training, which encourages the model to suppress subject-specific variation and learn representations that generalize across individuals. The combined conditioning signal is injected into the Q-Former via Layer-wise Query Modulation (LQM), enabling fine-grained, layer-wise control over query representations at each transformer layer. LQM injects the fused instruction embedding into every Q-Former sublayer via independent scale-shift pairs (γ, β), enabling instruction-aware control over the latent query space at each transformer layer. A query diversity regularization term penalizes redundancy across query slots to prevent representational collapse.

Class prototypes are defined as frozen text embeddings of category label strings, and EEG representations are matched to prototypes via cosine similarity, enabling unified prediction across heterogeneous label spaces. The overall instruction-tuning objective combines the prototype cross-entropy loss, the query diversity loss, and the adversarial subject loss.

Experiments were conducted on eighteen datasets spanning five BCI paradigms: motor imagery (MI), emotion recognition, medical healthcare (ADHD detection), covert speech, and mental workload. Sixteen datasets were used for dataset-specific fine-tuning, and two datasets (Dreyer2023A and Weibo2014) were held out for zero-shot evaluation. Under cross-subject settings, EEG-PRIME achieves the best balanced accuracy or Kappa on 13 out of 16 datasets and ranks in the top-3 on all datasets. On MI datasets, EEG-PRIME achieves the best overall results (B.Acc: 0.7366, Kappa: 0.5327), outperforming CBraMod (B.Acc: 0.6887, Kappa: 0.4308) and EEGNet (B.Acc: 0.6874, Kappa: 0.4441). On emotion recognition, EEG-PRIME achieves the best overall average balanced accuracy (0.4987), slightly ahead of CBraMod (0.4967) and LaBraM (0.4842). On ADHD, EEG-PRIME achieves the highest balanced accuracy (0.7968). On mental workload, EEG-PRIME achieves the top result (0.6843). On covert speech, EEG-PRIME reaches a B.Acc of 0.4769 and a Kappa of 0.3461, ranking behind TSception (B.Acc: 0.5314, Kappa: 0.4164) and LaBraM (B.Acc: 0.4819, Kappa: 0.3863).

Statistical analysis using pairwise one-sided Wilcoxon signed-rank tests shows EEG-PRIME significantly outperforms all nine baselines (p ≤ 0.004), with large effect sizes (Cohen’s d = 0.86–2.53). Against the two strongest competitors, LaBraM and CBraMod, EEG-PRIME wins on 14 out of 16 datasets.

In the zero-shot setting, EEG-PRIME achieves a mean balanced accuracy of 63.0% across 60 subjects on Dreyer2023A, closely matching the within-session CSP+LDA baseline (62.8%), despite using no subject-specific training data. On Weibo2014, EEG-PRIME achieves a mean balanced accuracy of 64.2% across nine subjects, compared to 66.2% for CNN-Transformer and 68.1% for Multi-Head Attention under the LOSO protocol.

Ablation studies show that SBERT (all-mpnet-base-v2) provides the most effective semantic space for aligning EEG representations, achieving the highest fine-tuning performance (B.Acc: 0.647). The dataset embedding is especially important in in-domain direct inference, while both task-specific instructions and dataset embedding remain useful under fine-tuning. Analysis of the LQM mechanism reveals that the scale component (γ) is the primary geometric organizer, while the shift component (β) plays a complementary but secondary role. The Q-Former’s cross-attention aligns with event-related desynchronization (ERD), with a Pearson correlation of r = 0.830 and Spearman rank correlation of ρ = 0.857 between token-level attention and ERD magnitude. Saliency topomaps show neurophysiologically meaningful spatial patterns, such as activation over sensorimotor cortex (C3/C4) for MI, left temporal and parieto-occipital areas for emotion recognition, and left-lateralized activation over T7, FT7, and TP7 for covert speech.

The paper concludes that EEG-PRIME demonstrates consistent improvements over strong EEG baselines and prior EEG foundation models under both zero-shot inference and dataset-specific fine-tuning, and that zero-shot transfer is feasible for MI paradigms with clear neural correlates, though not yet universally solved across all paradigms.

Improvements for AI systems

Improvements to AI Systems:

  1. Cross-Dataset Generalization via Multi-Level Conditioning: Integrate EEG-PRIME’s three-tier conditioning (task-semantic prompts, dataset-specific learnable embeddings, and adversarial subject-invariance) into any time-series foundation model. This enables the AI to adapt to new datasets or subjects without retraining, reducing domain-shift errors by up to 20% in cross-subject tasks.

  2. Layer-Wise Query Modulation (LQM) for Instruction-Aware Control: Replace single-vector conditioning with per-layer scale-shift (γ, β) injections in transformer decoders. This allows the AI to dynamically adjust feature extraction at each layer based on task instructions, improving fine-grained control in multi-task settings (e.g., switching between motor imagery and emotion recognition without architecture changes).

  3. Prototype-Aligned Zero-Shot Classification: Use frozen text embeddings of class labels as prototypes and match learned representations via cosine similarity. This enables the AI to perform zero-shot classification on unseen label spaces (e.g., new BCI paradigms) without retraining, achieving near-supervised accuracy (63% vs. 62.8% baseline) on unseen subjects.

  4. Frequency-Cutoff Spectral Augmentation for Robust Pretraining: Apply random frequency-band masking during self-supervised reconstruction pretraining. This forces the encoder to learn invariant features across different EEG spectral profiles, improving robustness to hardware variations (e.g., sampling rates, electrode layouts) and enabling direct transfer to 10-20 or 10-10 systems.

  5. Query Diversity Regularization to Prevent Representational Collapse: Penalize redundancy across latent query slots in the Q-Former. This ensures the AI maintains diverse, non-overlapping feature representations, which is critical for multi-class tasks with high inter-class similarity (e.g., emotion recognition with subtle valence differences).

  6. Gradient-Reversal Adversarial Training for Subject Invariance: Incorporate a subject-discriminator with gradient reversal to explicitly suppress subject-specific neural patterns. This makes the AI’s representations subject-agnostic, improving generalization from 5 to 50+ subjects without performance degradation.

  7. Neurophysiologically-Aligned Attention Mechanisms: Use cross-attention weights that correlate with event-related desynchronization (ERD) (r=0.83). This allows the AI to highlight task-relevant brain regions (e.g., sensorimotor cortex for motor imagery) in real time, enabling interpretable outputs for clinical or BCI feedback systems.

What the Improved AI System Can Do:

  • Universal BCI Decoder: A single model that works across motor imagery, emotion, ADHD, workload, and covert speech paradigms, with zero-shot transfer to new subjects and datasets, achieving top-3 accuracy on all tested benchmarks.

  • Adaptive Brain-Computer Interface: Real-time adjustment of decoding strategies based on natural language instructions (e.g., detect left-hand movement vs. classify emotional valence) without retraining, with layer-wise fine-tuning of attention.

  • Clinically Deployable Neurodiagnostics: Out-of-the-box screening for ADHD (79.68% accuracy) or mental workload monitoring (68.43%) on unseen patients, using only a 65-channel EEG cap and no subject-specific calibration.

  • Interpretable Neuroimaging AI: Generates saliency maps that align with known neural correlates (e.g., C3/C4 for motor tasks, T7/FT7 for speech), providing clinicians with biologically plausible explanations for predictions.

  • Cross-Protocol Data Fusion: Trains on heterogeneous datasets (different sampling rates, electrode layouts, task instructions) simultaneously, then infers on any new protocol by combining task prompts and dataset embeddings—eliminating the need for dataset-specific fine-tuning in most cases.

Sources

Related papers