Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay
summary
The gist
Understanding how humans and artificial intelligence systems predict and plan by interacting with their environment is a fundamental challenge at the intersection of neuroscience and machine learning.
In short
The study compared vision-language models (VLMs) and large-action models (LAMs) against reinforcement learning baselines during video games. Both outperformed RL baselines in brain encoding, showing that action fine-tuning reorganizes neural representations toward motor control despite similar overall prediction accuracy. Prompting gains are strongest in higher-order frontal regions, and LAMs show a distinct action-asymmetric organization.
Key concepts
- Voxel-wise Encoding Performance
- This measures how well the brain activity patterns of an AI model match the actual neural activity in specific brain regions (voxels). The study found that both VLMs and LAMs showed significantly better voxel-wise encoding than standard reinforcement learning methods, suggesting they learn more action-relevant representations.
- Prompt-Driven Gains
- This examines how providing specific instructions (prompts) changes the AI's internal brain representation. The largest improvements occurred in higher-order areas like frontal and parietal regions, indicating that prompting recruits the same brain networks used for planning in traditional reinforcement learning.
- Variance Partitioning
- This technique looks at how different types of information (like action versus reasoning) are organized within the model's representation. The finding showed that VLMs are prompt-symmetric, while LAMs are action-asymmetric, meaning action fine-tuning captures more unique variance than reasoning prompts.
Terminology used across episodes
This episode discusses
- Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay · Paper Radio
- Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners
- Vision-Language Integration in Multimodal Video Transformers (Partially) Aligns with the Brain
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- Human-Level Reinforcement Learning through Theory-Based Modeling, Exploration, and Planning
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- Atari-GPT: Benchmarking Multimodal Large Language Models as Low-Level Policies in Atari Games
- OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
The paper
Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay · Read on arXiv
Microsoft Research Bangalore, India · AWS AI Labs Amazon, USA
Understanding how humans and artificial intelligence systems predict and plan by interacting with their environment is a fundamental challenge at the intersection of neuroscience and machine learning. Most brain-encoding studies focus on aligning artificial models with brain activity during language comprehension or passive visual processing, while interactive brain alignment studies have to date been largely limited to reinforcement-learning (RL) agents and theory-based models. To address this gap, we study brain alignment of representative models from two foundation-model types, namely vision-language models (VLMs) and large-action models (LAMs), using fMRI recordings from participants playing naturalistic Atari-style video games. Specifically, we examine how action-focused and reasoning-focused prompts shape the models' internal representations and their alignment with fMRI brain activity. First, we find that both VLMs and LAMs achieve significantly higher voxel-wise encoding performance than RL baselines, with the advantage holding even under matched feature dimensionality. Second, compared to a no-prompt baseline, prompt-driven gains are larger in higher-order frontal-parietal and motor-planning regions than in early visual cortex, roughly 2-2.5 times when averaged over region-of-interest (ROI) groups, although individual regions are heterogeneous. Third, variance partitioning reveals a qualitatively different representational organization. VLM representations are prompt-symmetric (12.4% unique action vs. 9.5% unique reasoning), whereas LAM representations are action-dominant (25.6% unique action vs.-8.2% unique reasoning), with the asymmetry strongest in frontal-motor cortex. Together, these results associate action specialization with distinct cortical alignment patterns in multimodal game-state representations, revealing differences hidden by similar prediction accuracy.
Transcript
Introduction to the show: ident: Genomics Radio. Generated commentary on the latest computational biology and genomics papers.
Ines: I'm Ines, and with me are Marcus and Yuki, guest researcher.
Marcus: Today's paper: "Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay".
Ines: Understanding how humans and artificial intelligence systems predict and plan by interacting with their environment is a fundamental challenge at the intersection of neuroscience and machine learning.
Marcus: First, who's behind it and why it matters.
Paper summary: Ines: So, to summarize what this paper is about, it investigates how humans predict and plan by interacting with their environment by looking at brain alignment between vision-language models and large-action models while participants played Atari-style video games during fMRI.
Marcus: The central thesis is that they are testing whether internal representations in these AI systems more closely capture the world-modeling computations that underpin human adaptive behavior, contrasting them with older reinforcement learning agents and theory-based models which primarily focus on reward-based action learning.
Yuki: From a population genetics view, this is about seeing if the computational strategies employed by these modern AI architectures reflect fundamental principles of how biological systems manage complex interactions in their environment.
Ines: The study specifically looks at two foundation model families, vision-language models and large-action models, and they measure brain alignment using Pearson correlation between predicted fMRI activity from models trained on per-model per-prompt representations and the actual brain responses recorded during gameplay frames.
Marcus: They compare these against EMPA and D-DQN baselines, finding that both VLMs and LAMs significantly outperform those RL baselines in voxel-wise brain encoding, even when the feature dimensionality is matched to eight dimensions.
Yuki: That finding is significant because it suggests that these foundation models are capturing a richer structure of the world than purely reward-optimized agents do, which could be related to how different species have evolved their cognitive flexibility.
Ines: Furthermore, they examined prompt conditions, observing that prompt-driven gains scale with the cortical processing hierarchy, showing the largest improvements in frontal-parietal and motor-planning regions like MFG and SMA.
Marcus: That hierarchical scaling is important because it gives us a map of where the neural machinery is being recruited when these models are doing their best work in aligning with human planning activity.
Yuki: It connects the computational findings back to neurobiology by showing that the AI’s successful prediction isn't just about raw data processing, but about engaging specific, higher-order cognitive regions associated with planning.
Ines: And they also found a difference in how reasoning traces compare to action outputs, noting that generated reasoning-trace representations are substantially less brain-aligned than the final action output representations in models like Qwen3 point 5 <ref:2605.19352#pg0>.
Marcus: That asymmetry is a crucial piece of data because it suggests a separation between the representation used for generating an intended plan and the representation that actually maps onto observed motor control, which is something we need to keep tracking in our statistical analysis of these models.
Yuki: It helps us understand if the internal logic of an AI system aligns with the sequential nature of biological planning or if there’s a disconnect between high-level thought and low-level execution.
Ines: So, to put it plainly, this paper is mapping out where the computational structure of modern foundation models overlaps with the neural structures used for action planning and decision-making in humans during interactive tasks.
Conclusion: Ines: Thinking about the title, "Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay," it really suggests they are trying to understand the neural bridge between high-level thought processes and physical movement through these AI architectures.
Marcus: And that bridge is what they’re actually quantifying by showing how different model types, like VLMs versus LAMs, organize their internal data when they are processing real-time visual game frames and text prompts.
Yuki: For the wider implications, this research suggests that we can use these complex AI systems as a tool to probe the specific neural circuits that handle planning and action selection in ways that might be inaccessible through traditional neuroimaging alone.
Ines: What this means in simpler terms is that we are getting a better way to understand the biological mechanisms of planning by seeing how well these sophisticated models mimic or differ from human brain activity during these complex tasks.
Marcus: It points toward a future where we can develop AI systems that aren't just good at predicting an outcome, but systems whose internal workings are structurally aligned with the actual computational pathways used by biological organisms for decision-making and behavior.
Yuki: That could provide insights into how cognitive structures have been shaped through evolutionary pressures, showing us which organizational patterns in AI mirror the adaptive strategies we see in living species.
Ines: It really moves the conversation from just observing brain activity to understanding the computational structures that generate that activity, whether it’s in a human or an AI system.
Marcus: So, ultimately, this paper contributes to understanding how predictive models function by linking them directly to specific cortical regions responsible for motor control and policy-relevant computations during active interaction with an environment.
Yuki: It’s a valuable piece of work because it shows the potential for computational modeling to inform our fundamental theories about neural computation and adaptive behavior in living systems.
Ines: And that’s where we leave it for now, connecting the findings on model structure to the broader implications for understanding cognitive biology and AI development.
More episodes
- 2607.15989-Diffusion-induced instabilities promote cooperation in eco-evolutionary networks
- 2609.08081-Reliability assessment and multicenter clinical application of magnetic resonance methods for knee cartilage quantification
- 2502.17449-Non-Markovain Quantum State Diffusion for the Tunneling in SARS-COVID-19 virus
- 2512.10515-UNAAGI: Atom-Level Diffusion for Generating Non-Canonical Amino Acid Substitutions
- 2607.16479-The Site Frequency Spectrum in an Exponentially Growing Population with Selection
- 2501.07440-Attention when you need
- 2511.03503-Beta frequency shifts in decision making: Spectral fingerprints or communication channels?
- 2606.13017-Deep Sleep Classification via EEG Signal Criticality: A Passive BCI Approach for Sleep-Improvement Neurofeedback
- 2508.09037-Drivers of periodicity in population dynamic models of long-lived, large mammals
- 2512.17988-easyplater: The easy way to generate microplate designs deconvolved from multivariate clinical data