Scaling Vision Transformers for Functional MRI with Flat Maps
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Scaling Vision Transformers for Functional MRI with Flat Maps".
Jane: Scaling Vision Transformers for Functional MRI with Flat Maps introduces CortexMAE,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: Welcome back everyone! We're talking about a really interesting piece of research out there today on arXiv, titled "Scaling Vision Transformers for Functional MRI with Flat Maps." It looks like this paper is diving into how we can adapt powerful Vision Transformer models to handle functional MRI data.
Jane: That sounds fascinating, Tom. It seems like the main idea is that they're exploring a new way to represent fMRI signals so that AI models can actually learn from them effectively. I've read the summary, and it suggests they are looking at whether this adaptation can open up new applications in neuroscience by testing different data representations.
Lu: I find the concept of using intermediate data representations really intriguing; it’s like finding a sweet spot between two extremes, which is exactly what they are hypothesizing with their flat map approach. This idea that a representation might be better than the raw input or some simple pre-processing is a huge area for creative AI exploration.
Meng: From an engineering standpoint, I'm curious how this flat map projection actually works in practice; does it add too much computational overhead when we scale up these models? I need to understand the practical constraints before we get too excited about the theoretical potential.
Lalam: Considering all the factors, Lalam thinks that this work could significantly improve how we process complex biological data across different modalities because it finds a representation that balances complexity and utility.
Tom: Exactly, Meng! It’s not just about computation; it’s about finding the right structural bias to help the AI learn faster and better from the fMRI data. The authors are introducing a new model family called CortexMAE trained on 2 point 1K hours of open fMRI data from the Human Connectome Project.
Jane: So, what they claim is that this flat map projection allows them to adapt Vision Transformers to fMRI by converting three dee volumes into 2D maps, which are then treated like standard images for the ViT. It’s a clever way to inject some brain geometry information directly into the input representation.
Lu: That adaptation is what makes it compatible with the standard Vision Transformer architecture, which is a key part of their contribution, and it's interesting how this geometric bias influences the model's learning path. I’m excited to see how they leverage that structural information during the masked autoencoder training process.
Paper summary: Meng: I'm still thinking about the comparison with other inputs; they test flat maps against parcellation and volume-based embeddings, which suggests they are systematically exploring different trade-offs in dimensionality reduction. What I really want to know is how much performance gain we actually see when moving from one representation to the other in real-world scenarios.
Lalam: From my perspective, the finding that flat maps generally perform best for cognitive state decoding is very impactful because it points toward a more universally applicable method for understanding brain states.
Tom: That's exactly what they claim—that flat maps offer the best performance for state decoding, while also providing that "goldilocks zone" between lossy and dense representations, which is a major win for practical application. It’s about finding that perfect balance.
Jane: And they’ve even released an open evaluation suite called BRAINMARKS to make sure other researchers can test these models on seven public source datasets, which is a huge step for reproducibility in this field. It shows they aren't just doing isolated experiments.
Lu: The systematic scaling analysis they performed, observing "strict power law scaling" with dataset size, gives us some concrete evidence about how these types of models behave as we feed them more data. That kind of mathematical relationship is something we need to keep watching.
Meng: Scaling laws are important, but I also noted in the paper that the reconstruction loss improves with model size but then "saturates at depth nine (37M encoder parameters)," which tells me we might not need infinitely large models to get better performance on this specific task. That’s good news for deployability.
Lalam: And the compute analysis shows that the parcel MAE is two point five times faster to train than the flat map MAE, which in turn is one point eight times faster than volume-based models. That comparison gives us a clear picture of where the practical engineering benefits lie for training these foundation models.
Tom: So, to wrap up this initial look at "Scaling Vision Transformers for Functional MRI with Flat Maps," we see they've introduced a model that leverages geometric priors from the flat map projection to achieve top performance on cognitive state decoding.
Paper summary: Jane: And they’ve highlighted that this representation sits in a useful middle ground between simpler and more complex data handling methods, providing a good balance for their use in neuroscience research.
Lu: The implication here is that we might see these kinds of geometrically-informed adaptations applied to other complex time-series data where spatial structure matters, expanding the reach of Vision Transformer applications beyond pure vision tasks.
Meng: It seems like the next step for practical impact is building robust pipelines that can handle these 4D fMRI inputs efficiently, especially since they mentioned that input normalization is critical for state decoding because removing coordinate normalization causes a "dramatic loss of performance on state prediction". That's a practical hurdle for implementation.
Lalam: That dependency on specific normalization procedures highlights how sensitive these models are to their input preparation, which is something engineers need to keep in mind when deploying any foundation model like CortexMAE.
Tom: Absolutely, and the conclusion they draw is that CortexMAE demonstrates state-of-the-art performance on cognitive state decoding, even though they found a null result for subject-level trait prediction where no single model achieved clear state-of-the-art performance.
Jane: That distinction is important because it suggests that the dynamic nature of the fMRI signal might be better captured by this representation than static traits like subject characteristics. It points toward a future where AI models can better track dynamic brain states in real-time.
Lu: I think the broader implication is that these scaling laws and representation comparisons pave the way for developing more effective intermediate representations across various scientific domains, not just neuroscience. We're looking at a general methodology here.
Meng: For our work at the startup, the practical impact is seeing how much we can reduce model size while maintaining high performance on these types of complex, multimodal inputs, especially since they showed scaling limits.
Lalam: I feel that this research really has the potential to improve our understanding of how AI systems learn from complex biological data by providing a concrete framework for representation selection.
Tom: That’s all we have time for today on "Scaling Vision Transformers for Functional MRI with Flat Maps," but I hope you all found this deep dive into the paper really insightful!
Conclusion: Tom: So, we've been deep in the weeds looking at how this paper tackles scaling Vision Transformers for functional MRI using flat maps.
Jane: Exactly, and today we’re wrapping up by looking at what this whole endeavor means for us in a broader sense. The title itself, "Scaling Vision Transformers for Functional MRI with Flat Maps," really tells you that they're trying to make these powerful image models work with brain data in a way that allows them to handle huge amounts of information efficiently.
Lu: What I find fascinating is the core idea they propose here; it's about finding this middle ground representation between simple, low-resolution views and incredibly detailed voxel-by-voxel data.
Meng: From an engineering standpoint, it’s interesting how they manage to fit such a complex geometric constraint—the flat map projection—into a standard ViT architecture without completely breaking the established training procedures.
Lalam: And that efficiency is what makes this so compelling because it suggests we can build better AI tools for neuroscience by using representations that balance detail and computational cost beautifully.
Tom: Precisely, Lalam! They’re showing us a way to get top performance on decoding cognitive states while keeping the model size manageable, which is something engineers have been chasing for a long time.
Jane: It really boils down to how this paper suggests we can unlock new applications in understanding brain function by using these learned projections.
Lu: The implication for the field is that we might see this geometric approach applied to other complex time-series data where spatial structure plays a major role in interpretation.
Meng: I'm curious if this means we need to rethink how we preprocess multimodal brain data before feeding it into deep learning models.
Tom: That’s the big question, Meng! It opens up a whole new avenue for how we structure and process biological signals, and next up, we’ll look closely at the results regarding which specific representation actually wins out for different types of tasks.
cs.CV, cs.AI, q-bio.NC
Submitted: 2025-10-15
Updated: 2026-09-28
Code: https://github.com/MedARC-AI/CortexMAE
Importance score: 78/100
The gist: Scaling Vision Transformers for Functional MRI with Flat Maps introduces CortexMAE, a family of self-supervised foundation models trained on functional MRI (fMRI) data using a novel cortical flat map
Key concepts
- Flat Map Patch Embedding
- This technique converts 3D fMRI volumes into 2D activity maps using a cortical flat map projection. This allows the data to be treated like standard images for Vision Transformers, injecting brain geometry information directly into the model's input patches.
- Parcellation Patch Embedding
- Instead of using flat maps, this method treats each distinct brain region (parcel) as an independent time series. Patches are extracted from these time series, focusing on temporal dynamics within specific anatomical areas.
- Cognitive State Decoding
- This is the task where the model attempts to predict a subject's current mental state based on their fMRI activity. The flat map representation was specifically shown to perform best for this task, indicating it captures the necessary spatial and temporal features for understanding brain function.
Terminology
Summary
Scaling Vision Transformers for Functional MRI with Flat Maps introduces CortexMAE, a family of self-supervised foundation models trained on functional MRI (fMRI) data using a novel cortical flat map projection representation. This work matters because it explores whether adapting Vision Transformers to fMRI can unlock new applications in neuroscience by testing intermediate data representations, specifically demonstrating that the flat map representation offers the best performance for cognitive state decoding while providing a goldilocks zone
between lossy and dense representations.
The gist: The flat map representation of 3D fMRI volumes is hypothesized to be a goldilocks zone
of intermediate fMRI representations that effectively balances the extremes of parcellation and volume-based representations, performing best for state decoding.
Model Architecture and Representation
The core innovation is adapting the Vision Transformer (ViT) to fMRI by first converting each 3D fMRI volume into a 2D map using a cortical flat map projection. This representation maintains the full cortical fMRI signal while reducing dimensionality and injecting inductive bias from brain geometry, making it directly compatible with standard ViT patch embedding. The model family introduced is CortexMAE, a spatiotemporal masked autoencoder (MAE-st) trained on 2.1K hours of open fMRI data from the Human Connectome Project (HCP-YA).
Input Representations Compared
The study systematically compares three different input representations for CortexMAE:
-
Flat Map Patch Embedding: This involves converting 3D fMRI volumes into 2D activity flat maps, which are then resampled to a regular image grid, allowing the use of standard spacetime ViT patch embedding.
-
Parcellation Patch Embedding: This approach embeds each parcel time series independently using a time-only patch size, such as the Schaefer-400 parcellation.
-
Volume Patch Embedding: This models 4D fMRI data using a 4D patch embedding, but it is restricted to only the 100K voxels of neurally active gray matter (Schaefer cortex mask) to reduce sequence length.
Performance on Downstream Tasks
The performance of the different representations varies significantly depending on the downstream task:
(Table 2 summary)
(For subject-level trait prediction, e.g., ABIDE, ADHD200, ADNI, PPMI)
The results show no reliable differences between input representations,
likely due to small sample sizes. However, for age prediction (HCP-A), the volume model reliably outperforms the other two models.
(For cognitive state decoding)
The flat map representation performs best for state decoding, while volume works best for age prediction, and parcellation is most efficient. The paper notes that Flat maps perform best for state decoding.
Scaling Laws and Compute Analysis
The research establishes the first systematic scaling analysis for fMRI, observing strict power law scaling
with dataset size. For example, the test loss decreases with increasing training dataset size according to a strict power law (Figure 7a). Scaling with model size also shows that reconstruction loss improves with model size, but this improvement saturates at depth 9 (37M encoder parameters).
Compute cost analysis shows that the parcel MAE is 2.5× faster to train than the flat map MAE, which in turn is 1.8× faster than volume.
Benchmark Suite: BRAINMARKS
To ensure reproducibility, the authors created BRAINMARKS, an open fMRI foundation model benchmark suite supporting all current models evaluated on seven public source datasets. This suite includes:
-
Subject-level trait prediction benchmarks (e.g., ABIDE for Autism).
-
Dynamic cognitive state decoding benchmarks (e.g., HCP-YA Task21 and NSD COCO24).
Conclusion and Limitations
The study concludes that CortexMAE demonstrates SotA performance on cognitive state decoding.
However, it reports a challenging null result for subject-level trait prediction: no single model achieves clear state-of-the-art performance,
as all models struggle to outperform a simple functional connectivity baseline. Key limitations include the lack of reproducible benchmarks and the observation that while dynamic state prediction is robust, trait prediction performance is inconsistent across datasets. Future work will focus on developing more effective intermediate representations and scaling pretraining beyond single-source datasets.
Ablation Findings
Ablation experiments revealed that tube masking prevents local interpolation across time,
and patch norm
or PC norm
did not yield clear benefits for reconstruction target performance in the flat map model. Furthermore, input normalization is critical for state decoding: removing coordinate normalization results in a dramatic loss of performance on state prediction,
suggesting models rely on static structural features for trait prediction but dynamics for state prediction.
Improvements for AI systems
As a fastidious researcher, I have analyzed this manuscript, Scaling Vision Transformers for Functional MRI with Flat Maps.
The core innovation lies in adapting Vision Transformer (ViT) architectures—specifically the Masked Autoencoder (MAE-st)—to functional MRI (fMRI) data by representing 3D volumes as 2D flat maps
derived from cortical projections.
Here are the specific, high-impact improvements and capabilities this research enables for AI systems:
),
-
Improve the generalizability and efficiency of foundation models across diverse fMRI modalities by introducing the
flat map
representation. This addresses the current limitation where existing models are restricted to parcellation or volume-based representations. -
Develop a scalable, self-supervised learning (SSL) framework for fMRI data that leverages large, open datasets (like HCP-YA) efficiently, achieving strict power law scaling with respect to both dataset size and model capacity.
-
Create a reproducible and rigorous benchmarking suite (
BRAINMARKS
) that allows for fair comparison of different fMRI foundation models across subject-level trait prediction and dynamic cognitive state decoding tasks. -
Enable robust, low-latency decoding of dynamic cognitive states (e.g., classifying mental states based on short fMRI clips) by utilizing the optimal representation (flat maps) identified in the study, leading to superior performance over prior models in this domain.
-
Provide a method for
denoising
fMRI time series by leveraging population priors learned during pretraining, allowing AI systems to recover underlying spatiotemporal dynamics from noisy BOLD signals.
The improved AI system (CortexMAE family) can perform the following specific tasks:
-
Predict an individual's clinical or demographic traits (e.g., diagnosis of Autism, ADHD, Alzheimer's Disease, Age, Sex) with high reliability across diverse patient populations by leveraging a robust intermediate representation that balances structural detail and computational efficiency.
-
Decode a subject’s current mental or cognitive state in real-time from short fMRI recordings (e.g., classifying which of 21 cognitive conditions the subject is in) with significantly higher accuracy than previous models, making it useful for neurofeedback or clinical monitoring.
-
Act as a powerful denoising tool for raw fMRI time series, effectively separating true underlying brain dynamics from noise by exploiting learned population priors derived from massive pretraining on general human connectome data.
-
Serve as a scalable foundation model that can be fine-tuned or adapted to new neuroimaging datasets, providing consistent performance scaling laws based on the size of the training data and the complexity (depth) of the ViT architecture used.
Sources
- A Cookbook of Self-Supervised Learning
- Perception Encoder: The best visual embeddings are not at the output of the network
- On the Opportunities and Risks of Foundation Models
- Brain-Semantoks: Learning Semantic Tokens of Brain Dynamics with a Self-Distilled Foundation Model
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Training Compute-Optimal Large Language Models
- Scaling Laws for Neural Language Models
- SGDR: Stochastic Gradient Descent with Warm Restarts
- Decoupled Weight Decay Regularization
- Towards a general-purpose foundation model for fMRI analysis
- In Pursuit of Pixel Supervision for Visual Pre-training
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models