A Language-Guided Multimodal Foundation Model for Zero-Shot and Multi-Task Brain Signal Analysis
cs.LG, cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 38 pages, 7 main figures and 20 supplementary figures; includes Supporting Information. Code: https://github.com/mingzhi-c/metis-brain-signal-foundation-model
Journal ref: Advanced Intelligent Systems, e70486 (2026)
DOI: 10.1002/aisy.70486
Code: https://github.com/mingzhi-c/metis-brain-signal-foundation-model
License: http://creativecommons.org/licenses/by/4.0/
The gist: Brain signal analysis is essential for both neuroscience research and clinical diagnostics, yet current approaches face critical limitations.
Terminology
Abstract
Brain signal analysis is essential for both neuroscience research and clinical diagnostics, yet current approaches face critical limitations. End-to-end models require task-specific retraining and exhibit limited generalization, while pre-trained models lack semantic depth and still depend on extensive fine-tuning. Meanwhile, general-purpose multimodal foundation models, though powerful in other domains, struggle to interpret brain signals due to representational misalignment and lack of domain knowledge. This study introduces a multimodal foundation model for zero-shot and multi-task brain signal analysis (METIS) through a unified language-signal alignment framework. METIS is pretrained on the largest and most diverse brain-signal corpus to date, comprising over 70,000 h of recordings from more than 11,000 subjects across 20 datasets. In a comprehensive zero-shot evaluation across 12 datasets, METIS outperformed the leading generalist model by over 20.9% in average accuracy. Remarkably, without any fine-tuning, METIS's performance matches or exceeds that of supervised, task-specific models. Furthermore, METIS demonstrates exceptional data efficiency and strong generalization, achieving an average AUROC advantage of over 16.0% in few-shot settings and 15.9% in cross-dataset transfer. This work establishes a new paradigm for general-purpose brain signal analysis, paving the way for next-generation neurotechnology.
Sources
- GPT-4 Technical Report
- Gemini: A Family of Highly Capable Multimodal Models
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- Qwen2.5 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks