Continuous-Latent Predictive Modeling with Semantic Alignment for EEG-Language Foundation Models
Myeong-Ju Cho, Hye-Bin Shin, Seo-Hyun Lee, Seong-Whan Lee
Korea University
cs.LG
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 19 pages, 3 figures; supplementary material included
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: The paper proposes the Brain Latent Predictive Model (BLPM), an EEG-language foundation model that reformulates heterogeneous EEG decoding tasks as a continuous semantic embedding prediction problem.
Terminology
Summary
The paper proposes the Brain Latent Predictive Model (BLPM), an EEG-language foundation model that reformulates heterogeneous EEG decoding tasks as a continuous semantic embedding prediction problem. The authors state: we propose Brain Latent Predictive Model (BLPM), an EEG–language foundation model that reformulates heterogeneous EEG decoding tasks as a continuous semantic embedding prediction problem.
The model addresses limitations of existing EEG foundation models. The paper notes: "dominant pretraining paradigms face key challenges: masked autoencoding tends to prioritize low-level signal reconstruction over task-relevant semantics, while autoregressive modeling creates a mismatch between continuous neural dynamics and discrete token spaces. The authors explain that
masked autoencoding tends to focus on reconstructing low-level signal structures and that
quantizing continuous neural dynamics into a finite codebook may discard discriminative patterns that are critical for distinguishing task-relevant neural states."
BLPM consists of three main components. First, the Continuous EEG Latent Predictive (CELP) encoder learns transferable representations through latent target prediction.
The paper describes: "Rather than directly reconstructing raw EEG waveforms, BLPM employs a Continuous EEG Latent Predictive (CELP) Encoder that predicts latent representations of target EEG segments from contextual EEG observations, thereby reducing excessive dependence on low-level waveform details and subject-specific variations while learning transferable representations that capture higher-level neurophysiological structures. The CELP encoder uses
a joint-embedding predictive architecture with an online encoder, a target encoder updated via exponential moving average, and an embedding predictor, trained with
a Smooth L1 loss over masked positions" in the latent space.
Second, the Multi-Query Semantic Decomposition (MQSD) module extracts task-relevant information and aligns continuous EEG representations with textual semantics within a shared latent space according to their semantic relationships.
The paper states: MQSD introduces a fixed set of M semantic queries
that correspond to task-relevant EEG factors, including temporal dynamics, spectral characteristics, spatial relationships, and morphological patterns.
Each query independently attends to the continuous EEG tokens through semantic query cross-attention,
and the resulting summaries are projected into the hidden dimension of the LLM.
Third, the model uses multi-task instruction tuning with semantic answer matching
for unified EEG decoding. The paper explains: Instead of autoregressively generating answer tokens, BLPM predicts the semantic embedding of the target answer in a single forward pass.
During inference, the predicted label is selected by semantic matching within the task-specific candidate set
using cosine similarity.
The main contributions listed in the paper are: proposing BLPM for unified EEG decoding without discrete neural tokenization or autoregressive generation
; introducing the CELP encoder whose latent predictive objective promotes higher-level abstraction without directly reconstructing low-level details
; presenting the MQSD module that decomposes continuous EEG representations into multiple task-relevant semantic queries with language-derived semantic guidance
; and systematically evaluating BLPM using standardized benchmarks to enable fair comparisons.
The model is pretrained on the Temple University Hospital EEG Corpus (TUEG), described as a large-scale publicly available clinical EEG dataset comprising 69,652 recordings from 14,987 subjects, with a total of 27,062 hours of recordings.
Downstream evaluation covers seven tasks: "mental workload classification (COG-BCI), mental stress detection (Mental Arithmetic), abnormal detection (TUAB), event type classification (TUEV), motor imagery classification (PhysioNet-MI), emotion recognition (FACED), and sleep staging (HMC). The evaluation uses the NeuralBench framework for
standardized evaluation protocols."
Results show BLPM outperforms baselines. The paper reports: BLPM demonstrates strong generalization capabilities across heterogeneous EEG decoding tasks.
Compared to task-specific models, BLPM consistently outperforms the task-specific baselines, EEGNet and EEGConformer, across all seven downstream tasks,
with performance improvements of 14.0%, 10.6%, and 9.9% over the best-performing task-specific baseline on Mental Arithmetic, TUEV, and PhysioNet-MI.
Compared to foundation models, BLPM achieves a higher balanced accuracy on six datasets,
including a balanced accuracy of 34.3% on FACED, 75.4% on Mental Arithmetic, and 57.2% on TUEV, outperforming REVE, the strongest EEG foundation model baseline, by 2.3%, 2.1%, and 1.9%, respectively.
BLPM achieves these results using a unified multi-task model, whereas the competing foundation models are separately fine-tuned on individual datasets.
Ablation studies confirm the contribution of each component. Replacing latent prediction with masked reconstruction reduces performance by up to 2.55%p.
Removing the alignment stage causes the largest decrease of 6.91% observed on FACED.
The MQSD module outperforms global pooling and generic learnable queries. Replacing semantic embedding prediction with autoregressive token prediction yields improvements of 1.91%, 1.37%, and 1.14%
in favor of the embedding prediction approach.
The paper concludes: "By replacing discrete tokenization and autoregressive modeling with continuous semantic prediction, BLPM aligns EEG representations with language semantics in a continuous latent space rather than a discrete token space. Across diverse EEG benchmarks, BLPM demonstrates strong performance on heterogeneous neural decoding tasks. These results support continuous latent prediction as an effective modeling paradigm for universal EEG decoding."
Improvements for AI systems
Improvements to AI systems:
-
Continuous semantic embedding prediction instead of discrete token generation – Replace autoregressive token prediction with direct prediction of semantic embeddings in a shared latent space. This avoids the information loss from quantizing continuous signals into discrete tokens and enables single-forward-pass inference, reducing latency and computational cost.
-
Latent predictive pretraining over masked reconstruction – Train encoders to predict latent representations of masked input segments (e.g., EEG windows) rather than reconstructing raw signals. This forces the model to learn high-level, task-relevant abstractions and reduces sensitivity to low-level noise and subject-specific variations.
-
Multi-query semantic decomposition for heterogeneous inputs – Use a fixed set of learned semantic queries that independently attend to different aspects of the input (e.g., temporal, spectral, spatial, morphological). This allows the model to extract diverse, task-relevant features and align them with language semantics, improving generalization across multiple downstream tasks without task-specific architectures.
-
Unified multi-task instruction tuning with semantic answer matching – Instead of generating free-form text, predict the semantic embedding of the correct answer and select the best match from a task-specific candidate set via cosine similarity. This enables a single model to handle many tasks (e.g., classification, detection, staging) without separate fine-tuning, improving scalability and cross-task transfer.
-
Joint-embedding predictive architecture with EMA target encoder – Use an online encoder and a target encoder updated via exponential moving average to stabilize training and avoid representation collapse, leading to more robust and transferable representations.
-
Language-guided alignment in continuous latent space – Align neural representations with textual semantics during pretraining and fine-tuning, enabling zero-shot or few-shot generalization to new tasks where textual descriptions are available.
What the improved AI system can do:
-
Process continuous, non-linguistic signals (e.g., EEG, physiological data, sensor streams) and map them to semantic embeddings that are directly comparable to language embeddings, enabling cross-modal reasoning and instruction-following.
-
Handle multiple heterogeneous tasks with a single unified model – e.g., classify mental workload, detect stress, recognize emotions, stage sleep, and identify abnormal events – without retraining per task, using only task-specific candidate labels.
-
Achieve higher accuracy on low-resource or noisy signal data by focusing on high-level semantics rather than raw signal reconstruction, making it robust to individual differences and sensor noise.
-
Perform fast, low-latency inference by predicting embeddings in one forward pass, suitable for real-time brain-computer interfaces or adaptive systems.
-
Transfer knowledge across tasks and domains – e.g., pretrain on large clinical EEG data, then adapt to new tasks with minimal labeled data by leveraging the shared semantic space.
-
Provide interpretable multi-aspect representations via the multi-query decomposition, allowing users to inspect which signal characteristics (temporal, spectral, spatial) drive predictions.
Sources
- NeuralBench: A Unifying Framework to Benchmark NeuroAI Models
- BrainRVQ: A High-Fidelity EEG Foundation Model via Dual-Domain Residual Quantization and Hierarchical Autoregression
- The Llama 3 Herd of Models
- NeuralSet: A High-Performing Python Package for Neuro-AI
- Neural Signals Generate Clinical Notes in the Wild
- EmbeddingGemma: Powerful and Lightweight Text Representations
- KAST-BAR: Knowledge-Anchored Semantically-Dynamic Topology Brain Autoregressive Modeling for Universal Neural Interpretation
- VP-VAE: Rethinking Vector Quantization via Adaptive Vector Perturbation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks