A foundation model with multi-variate parallel attention to generate neuronal activity
cs.LG, cs.AI
Submitted: 2025-06-25
Updated: 2026-08-31
Comments: The code is available at https://github.com/IBM/multi-variate-parallel-transformer. The SWEC iEEG dataset is available at https://huggingface.co/datasets/NeuroTec/SWEC_iEEG_Dataset. Published at ICLR 2026
Code: https://github.com/IBM/multi-variate-parallel-transformer
License: http://creativecommons.org/licenses/by/4.0/
The gist: Learning from multi-variate time-series with heterogeneous channel configurations remains a fundamental challenge for deep neural networks, particularly in clinical domains such as intracranial
Terminology
Abstract
Learning from multi-variate time-series with heterogeneous channel configurations remains a fundamental challenge for deep neural networks, particularly in clinical domains such as intracranial electroencephalography (iEEG), where channel setups vary widely across subjects. In this work, we introduce multi-variate parallel attention (MVPA), a novel self-attention mechanism that disentangles content, temporal, and spatial attention, enabling flexible, generalizable, and efficient modeling of time-series data with varying channel counts and configurations. We use MVPA to build MVPFormer, a generative foundation model for human electrophysiology, trained to predict the evolution of iEEG signals across subjects. To support this and future efforts by the community, we release the SWEC iEEG dataset, the largest publicly available iEEG dataset to date, comprising nearly 10,000 hours of recordings from heterogeneous clinical sources. MVPFormer leverages MVPA to achieve strong generalization across subjects, demonstrating expert-level performance in several iEEG tasks. MVPFormer surpasses state-of-the-art (SOTA) Transformer baselines in seizure detection across the SWEC, the MAYO, and the FNUSA datasets, while also achieving SOTA performance on four Brain TreeBank iEEG decoding tasks (volume, pitch, onset, and speech). We further validate MVPA on standard time-series forecasting and classification tasks, where it matches or exceeds the performance of existing attention-based models. Together, our contributions establish MVPA as a general-purpose attention mechanism for heterogeneous time-series and MVPFormer as the first open-source, open-weights, and open-data iEEG foundation model with SOTA clinical performance. The code is available at https://github.com/IBM/multi-variate-parallel-transformer. The SWEC iEEG dataset is available at https://huggingface.co/datasets/NeuroTec/SWEC iEEG Dataset.
Sources
- Graph-based Clustering for Detecting Semantic Change Across Time and Languages
- Transformers in Time Series: A Survey
- Generating Long Sequences with Sparse Transformers
- LLM Pretraining with Continuous Concepts
- Training Large Language Models to Reason in a Continuous Latent Space
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Conformer: Convolution-augmented Transformer for Speech Recognition
- Scaling Laws for Neural Language Models
- Neuro-GPT: Towards A Foundation Model for EEG
- Axial Attention in Multidimensional Transformers
- Representation Learning with Contrastive Predictive Coding
- Objective evaluation metrics for automatic classification of EEG events
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks