A Multimodal Sequence-to-Sequence Model for Cross-Subject Prediction of Brain Responses to Naturalistic Stimuli
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "A Multimodal Sequence-to-Sequence Model for Cross-Subject Prediction of Brain Responses to Naturalistic Stimuli".
Jane: The gist This sequence-to-sequence Transformer predicts whole-brain fMRI responses to naturalistic multimodal movies by autoregressively predicting activity from visual, auditory, and language inputs.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Now let's talk about the core idea behind this work, "A Multimodal Sequence-to-Sequence Model for Cross-Subject Prediction of Brain Responses to Naturalistic Stimuli," and who came up with it. The authors are Qianyi He and Yuan Chang Leong from the Data Science Institute at the University of Chicago.
Jane: They’re tackling a problem that moves beyond simple correlation, trying to model how different people react differently to the exact same movie or scene using this sequence modeling framework. It’s about capturing individual brain responses alongside general stimulus features.
Lu: The title itself points to their main contribution: cross-subject prediction of brain responses. They are explicitly trying to account for the fact that while the input stimulus is the same, the resulting neural activity changes depending on who is watching it.
Meng: So, they’re not just building a model that predicts *a* brain response; they are building one where the prediction is tailored to an individual subject, which I think makes their architecture much more sophisticated than what we usually see in these kinds of studies.
Lalam: It's about making the system robust enough so it can learn general patterns from all subjects while still being able to adapt quickly when it meets a new person.
The paper's summary: Tom: So, let’s look at what the authors actually say in their summary. They propose an autoregressive Transformer that predicts fMRI activity by taking these multimodal inputs—visual, auditory, and language—and feeding them into a decoder that builds the brain signal step by step.
Jane: The key is how they feed information into that decoder. They use dual cross-attention mechanisms to pull in both the perceptual details from the current stimulus and some high-level narrative context derived from summaries of the content, like episode descriptions.
Lu: They are using sequences of multimodal context to predict sequences of brain activity, which they say lets them capture long-range temporal structure in both the stimuli and the neural responses simultaneously. That’s a big claim about capturing that deep connection between input and output over time.
Meng: And they’re also doing this by using a shared encoder with subject-specific decoder heads, which is a way to make sure the general parts of the learning are shared across subjects, but the final mapping to brain activity is unique per person.
Lalam: They are using autoregressive decoding with teacher forcing and learnable BOS tokens to predict the entire sequence of fMRI data, which helps stabilize training and lets it generate these long sequences reliably.
The paper's improvements: Tom: When we look at what they suggest as improvements, they focus on three main areas. First is that autoregressive sequence generation for brain activity prediction, which simulates the way subjects actually watch things by predicting the future based on past inputs.
Jane: Second is this multimodal context integration via dual cross-attention, where the decoder attends to both the sensory input and those high-level narrative summaries we talked about earlier. That’s a smart way to give it context beyond just what’s happening right now.
Lu: And third, they tackle inter-subject variability with their hybrid architecture—the shared encoder and subject-specific decoder heads—which is designed to leverage common representations while still allowing for that individual adaptation you mentioned earlier.
Meng: From an engineering standpoint, the sliding window augmentation strategy is also a key part of making this work on real data; they segment stimuli into overlapping temporal chunks of forty time steps, and the output sequence length ends up being thirty-five frames after a hemodynamic delay.
Lalam: Plus, they use teacher forcing with learnable BOS tokens during training, which helps them stabilize the learning process and accelerates convergence when predicting these long sequences.
Conclusion: Tom: So to wrap up on this paper, the idea is that by using this multimodal sequence-to-sequence Transformer model for cross-subject prediction of brain responses to naturalistic stimuli, they can predict whole-brain fMRI activity from visual, auditory, and language inputs.
Jane: They’ve shown it can handle the complexity of these different modalities while keeping track of individual differences between subjects through that hybrid architecture setup. It proves that a shared encoder for general stimulus features combined with subject-specific heads is a viable path forward here.
Lu: The implication is that we can move toward models that capture the temporal dependencies in both complex sensory inputs and the resulting neural activity in a unified, generative way, which opens up new avenues for understanding perception.
Meng: Practically speaking, this shows how you can use these large sequence models to build predictive tools for neuroscience data where individual variability is a big factor. It gives us a framework to predict brain responses tailored to specific people.
Lalam: It means that the advances in these multimodal sequence models can help us create systems that are better at understanding and modeling complex human perception by bridging the gap between what we see or hear and what we actually feel in our brains.
Qianyi He, Yuan Chang Leong
Data Science Institute, University of Chicago · Department of Psychology, Neuroscience Institute, University of Chicago
cs.CV, q-bio.NC
Submitted: 2025-07-24
Updated: 2026-10-05
Comments: Substantially revised manuscript with new analyses, expanded cross-subject evaluation, and updated figures
Code: https://github.com/Angelneer926/Algonauts_challenge
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 79/100
The gist: The gist This sequence-to-sequence Transformer predicts whole-brain fMRI responses to naturalistic multimodal movies by autoregressively predicting activity from visual, auditory, and language inputs.
Key concepts
- VideoMAE
- This is a pre-trained model used to extract meaningful representations from video frames. It helps the system understand the temporal dynamics of visual motion, allowing it to create a richer understanding of what is happening visually at each point in time.
- HuBERT
- HuBERT is a self-supervised model applied to audio signals. It learns hierarchical speech representations by predicting masked parts of the audio, capturing both the specific sounds (phonetics) and the emotional tone or rhythm (prosody) of speech.
- Sequence-to-Sequence Transformer
- This is a type of deep learning architecture designed to map one sequence (the multimodal movie inputs: video, audio, text) to another sequence (the fMRI brain activity). It allows the model to generate the brain response step-by-step based on all the preceding inputs.
- Subject Variability Handling
- The model uses a shared encoder but has separate decoder heads for each subject. Each subject gets a unique embedding vector that is added to every time step of the decoder, allowing it to adapt its prediction specifically to that individual's brain response patterns.
Terminology
Summary
The gist
This sequence-to-sequence Transformer predicts whole-brain fMRI responses to naturalistic multimodal movies by autoregressively predicting activity from visual, auditory, and language inputs.
Feature Extraction
-
Visual Modality: The model employs VideoMAE to extract a 768-dimensional embedding for the current frame by averaging token representations from the final (12th) encoder layer, which is then concatenated with the mean embedding of all preceding frames, resulting in a 1536-dimensional vector
-
Audio Modality: HuBERT is utilized to extract a 1536-dimensional feature vector at each time point by concatenating the hidden states from the 3rd and 9th encoder layers, enabling the capture of both phonetic and prosodic information
-
Language Modality: The Qwen model embeds the dialogue transcript, receiving the current utterance along with all previous utterances as textual context, truncated to the most recent 2048 tokens
-
Visual/Language Hybrid: BridgeTower is used to extract a 1536-dimensional fused representation from the pooler output of the final layer, which corresponds to the [CLS] token after visual/text fusion
-
Narrative Summaries: Sentence-level embeddings from a pretrained BERT model are obtained for narrative summaries, and two strategies are explored to incorporate this additional context into frame-wise representation learning processes
Model Architecture
-
Encoder Design: The encoder is implemented as a Transformer with causal self-attention, such that each time step only attends to its current and preceding inputs, which is motivated by the nature of neural processing
-
Decoder Design: The decoder is a masked causal Transformer that generates fMRI signals autoregressively, where at each decoding step t, the decoder receives as query the embedding of the previously generated output yˆt−1 and attends only to the past via masked self-attention
-
Contextual Integration: To provide high-level semantic context, an additional cross-attention module over episode-level descriptions E is introduced after the standard cross-attention over stimulus features, allowing the decoder to attend to both perceptual and narrative information during fMRI prediction
-
Subject Variability Handling: A hybrid architecture with a shared encoder and subject-specific decoder heads is adopted, where each subject s is associated with a learnable embedding vector es that is concatenated to every time step of the decoder input
Training Strategy
-
Sliding Window Augmentation: The dataset is augmented using a sliding window approach, segmenting stimuli into overlapping temporal chunks of 40 consecutive time steps, with the corresponding output being a 35-frame fMRI sequence temporally aligned after applying a fixed hemodynamic delay
-
Loss Function: The training objective combines mean squared error (MSE) and negative Pearson correlation between the predicted and ground truth sequences, defined as LMSE = 1/T Sum(yˆt − yt2) + λLcorr where ρ denotes the Pearson correlation
-
Autoregressive Decoding Guidance: The model employs a learnable [BOS] token and teacher forcing strategy, where with probability γ, the ground truth is used to stabilize training and accelerate convergence
The model achieves strong performance on both in-distribution and out-of-distribution data, demonstrating the effectiveness of temporally-aware, multimodal sequence modeling for brain activity prediction. The final model achieved an average Pearson correlation of 0.305 on Friends Season 7 and 0.199 on the OOD movie set. This work presents a multimodal sequence-to-sequence Transformer model for predicting full-brain fMRI responses to naturalistic audiovisual and linguistic stimuli, capturing both the temporal and individual complexity of neural responses. The results support the promise of generative, multimodal models as a powerful framework for bridging perception and brain activity in real-world settings.
--- Page 1 ---
A MULTIMODAL SEQ2SEQ TRANSFORMER FOR PREDICTING BRAIN RESPONSES TO NATURALISTIC STIMULI Qianyi He Data Science Institute University of Chicago, USA heqianyi926@uchicago.edu Yuan Chang Leong Department of Psychology, Neuroscience Institute University of Chicago, USA ycleong@uchicago.edu ABSTRACT The Algonauts 2025 Challenge called on the community to develop encoding models that predict whole-brain fMRI responses to naturalistic multimodal movies In this submission, we propose a sequence-to-sequence Transformer that autoregressively predicts fMRI activity from visual, auditory, and language inputs Stimulus features were extracted using pretrained models including VideoMAE, HuBERT, Qwen, and BridgeTower The decoder integrates information from prior brain states and current stimuli via dual cross-attention mechanisms that attend to both perceptual information extracted from the stimulus as well as narrative information provided by high-level summaries of the content One core innovation of our approach is the use of sequences of multimodal context to predict sequences of brain activity enabling the model to capture long-range temporal structure in both stimuli and neural responses Another is the combination of a shared encoder with partial subjectspecific decoder, which leverages common representational structure across subjects while accounting for individual variability Our model achieves strong performance on both in-distribution and out-ofdistribution data demonstrating the effectiveness of temporally-aware, multimodal sequence modeling for brain activity prediction The code is available at https://github.com/Angelneer926/ Algonauts challenge Keywords neural encoding sequence-to-sequence transformer naturalistic movies fMRI
--- Page 2 ---
Traditional approaches to modeling brain responses such as finite impulse response (FIR) models or ridge regression typically learn linear mappings that predict each timepoint in the fMRI signal independently from recent stimulus history [2, 3] While these methods have been successful in many settings they often neglect the dynamic autoregressive nature of neural responses and are limited in their ability to integrate multimodal inputs or adapt to individual differences across subjects Inspired by recent studies demonstrating that deep neural networks can effectively model cortical responses to naturalistic language and vision [4, 5, 6] we propose a sequence-to-sequence Transformer-based architecture to better capture the temporal dependencies and multimodal interactions that shape fMRI signals over time Similar to neural network-based machine translation models that map sequences of words from one language to another [7] our model translates sequences of audiovisual and linguistic stimuli into sequences of brain responses conditioning each prediction on both the full input sequence and the history of prior neural activity To account for the rich, hierarchical structure of naturalistic stimuli we extracted stimulus features from state-of-the-art pretrained models across multiple modalities: VideoMAE for temporal dynamics of visual motion [8] acoustic features [9] and Qwen for linguistic representations Additionally we incorporated sentence-level semantic features derived from BERT [10] using high-level summaries of the TV episodes and movies that constitute the training and testing data This allows us to provide broader narrative context beyond the moment-to-moment stimulus input [11] To capture joint visual-linguistic representations we extracted cross-modal features from the stimuli using BridgeTower [12]
--- Page 3 ---
The Algonauts 2025 Challenge is based on a subset of the CNeuroMod dataset[25] and includes whole-brain fMRI recordings from four participants (sub-01, sub-02, sub-03, sub-05) while they viewed naturalistic multimodal movies The training data includes approximately 65 hours of movie stimuli from all episodes of Friends Seasons 1–6 and four feature films The Bourne Supremacy Hidden Figures Life The Wolf of Wall Street each aligned with time-resolved fMRI responses fMRI responses have been preprocessed and parcellated into 1,000 cortical regions using the Schaefer atlas [26] sampled every 1.5 seconds and aligned to MNI space Model performance is evaluated using the Pearson correlation between predicted and ground-truth fMRI responses The final score is obtained by averaging the correlation values across all subjects test movies and brain parcels
--- Page 4 ---
We extracted stimulus representations from multiple modalities to comprehensively capture the external input at each time point See Figure 1 This section outlines the models and strategies used to obtain and enhance these features Visual Modality For visual input we employed VideoMAE a masked autoencoder pretrained on large-scale video datasets VideoMAE is well-suited for this task due to its ability to model fine-grained spatiotemporal dependencies particularly motion dynamics and scene transitions At each time point we extracted a 768-dimensional embedding for the current frame by averaging token representations from the final (12th) encoder layer We then concatenated this with the mean embedding of all preceding frames resulting in a 1536-dimensional vector This simple yet effective temporal smoothing provides a form of memory that helps the model maintain awareness of past visual context which is especially beneficial for videos with continuous narrative flow
Figure 1: Feature Extraction
Audio Modality To encode the audio signal we utilized HuBERT a self-supervised model that learns hierarchical speech representations through masked prediction of latent units We extracted a 1536-dimensional feature vector at each time point by concatenating the hidden states from the 3rd and 9th encoder layers enabling the capture of both phonetic and prosodic information While we experimented with incorporating past context e.g.
Improvements for AI systems
- textbf Autoregressive Sequence Generation for Brain Activity Prediction
The model is improved to generate fMRI time series autoregressively,
which simulates the temporal unidirectionality of how subjects viewed the stimuli.
This allows the AI system to predict future brain states conditioned on prior inputs, capturing dynamic neural responses rather than static mappings.
- textbf Multimodal Context Integration via Dual Cross-Attention
The decoder integrates information from "prior brain states and current stimuli via dual cross-attention mechanisms that attend to both perceptual information extracted from the stimulus as well as narrative information provided by high-level summaries of the content." This enables the system to leverage both moment-to-moment sensory data and overarching story context for richer predictions.
- textbf Subject-Specific Adaptation for Inter-Subject Variability
The architecture employs a hybrid architecture with a shared encoder and subject-specific decoder heads,
where each subject has a learnable embedding vector es that is concatenated to every time step of the decoder input.
This allows the system to predict fMRI responses tailored specifically to an individual's unique brain response patterns while benefiting from shared, generalizable stimulus representations.
- textbf Robust Training via Sliding Window and Teacher Forcing
To address data scarcity, a sliding window approach
segments stimuli into overlapping temporal chunks,
increasing training samples. Furthermore, the system uses teacher forcing where With probability γ, the ground truth is used to stabilize training and accelerate convergence; with probability 1 − γ, the model samples from its own output distribution.
This strategy ensures stable learning and robustness during inference.
- textbf Enhanced Feature Extraction for Temporal Dynamics
The visual modality uses a concatenation strategy: We then concatenated this [frame embedding] with the mean embedding of all preceding frames, resulting in a 1536-dimensional vector.
This provides a form of memory that helps the model maintain awareness of past visual context,
improving the capture of continuous motion dynamics.
Abstract
Brain encoding models predict time-resolved neural activity from computational representations of ongoing experience, providing a principled framework for testing how information is represented and transformed across cortical systems. Naturalistic audiovisual narratives are a particularly rich but challenging testbed for these models, requiring integration of multimodal inputs over long temporal horizons and generalization across individuals with substantial response variability. We introduce a multimodal sequence-to-sequence Transformer with a hybrid cross-subject parameterization that predicts cortex-wide parcel-wise fMRI time series autoregressively from visual, audio, language, and vision--language representations. We evaluate the approach on data from the Courtois NeuroMod project, where four deeply-sampled participants viewed six seasons of Friends and four feature-length films during fMRI. Sequence-to-sequence temporal modeling yields consistent improvements over single-frame prediction across cortical networks, with gains extending to novel stimuli. A hybrid architecture that pairs a shared stimulus encoder with lightweight subject-specific decoder components outperforms both fully shared and fully individual models, indicating complementary advantages of learning shared stimulus representations across subjects and fitting individual neural readouts. Finally, we show that in data-scarce settings, hybrid models can be personalized to new individuals with limited fMRI data, demonstrating that multi-subject pretraining serves as a strong inductive prior for building individual-specific encoding models. Together, these results indicate that combining multimodal sequence modeling with a hybrid cross-subject architecture offers a scalable framework for personalized brain encoding under naturalistic conditions.
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models