Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Genomics Radio. Generated commentary on the latest computational biology and genomics papers.
Ines: I'm Ines, and with me are Marcus and Yuki, guest researcher.
Marcus: Today's paper: "Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay".
Ines: Understanding how humans and artificial intelligence systems predict and plan by interacting with their environment is a fundamental challenge at the intersection of neuroscience and machine learning.
Marcus: First, who's behind it and why it matters.
Paper summary: Ines: So, to summarize what this paper is about, it investigates how humans predict and plan by interacting with their environment by looking at brain alignment between vision-language models and large-action models while participants played Atari-style video games during fMRI.
Marcus: The central thesis is that they are testing whether internal representations in these AI systems more closely capture the world-modeling computations that underpin human adaptive behavior, contrasting them with older reinforcement learning agents and theory-based models which primarily focus on reward-based action learning.
Yuki: From a population genetics view, this is about seeing if the computational strategies employed by these modern AI architectures reflect fundamental principles of how biological systems manage complex interactions in their environment.
Ines: The study specifically looks at two foundation model families, vision-language models and large-action models, and they measure brain alignment using Pearson correlation between predicted fMRI activity from models trained on per-model per-prompt representations and the actual brain responses recorded during gameplay frames.
Marcus: They compare these against EMPA and D-DQN baselines, finding that both VLMs and LAMs significantly outperform those RL baselines in voxel-wise brain encoding, even when the feature dimensionality is matched to eight dimensions.
Yuki: That finding is significant because it suggests that these foundation models are capturing a richer structure of the world than purely reward-optimized agents do, which could be related to how different species have evolved their cognitive flexibility.
Ines: Furthermore, they examined prompt conditions, observing that prompt-driven gains scale with the cortical processing hierarchy, showing the largest improvements in frontal-parietal and motor-planning regions like MFG and SMA.
Marcus: That hierarchical scaling is important because it gives us a map of where the neural machinery is being recruited when these models are doing their best work in aligning with human planning activity.
Yuki: It connects the computational findings back to neurobiology by showing that the AI’s successful prediction isn't just about raw data processing, but about engaging specific, higher-order cognitive regions associated with planning.
Ines: And they also found a difference in how reasoning traces compare to action outputs, noting that generated reasoning-trace representations are substantially less brain-aligned than the final action output representations in models like Qwen3 point 5 <ref:2605.19352#pg0>.
Marcus: That asymmetry is a crucial piece of data because it suggests a separation between the representation used for generating an intended plan and the representation that actually maps onto observed motor control, which is something we need to keep tracking in our statistical analysis of these models.
Yuki: It helps us understand if the internal logic of an AI system aligns with the sequential nature of biological planning or if there’s a disconnect between high-level thought and low-level execution.
Ines: So, to put it plainly, this paper is mapping out where the computational structure of modern foundation models overlaps with the neural structures used for action planning and decision-making in humans during interactive tasks.
Conclusion: Ines: Thinking about the title, "Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay," it really suggests they are trying to understand the neural bridge between high-level thought processes and physical movement through these AI architectures.
Marcus: And that bridge is what they’re actually quantifying by showing how different model types, like VLMs versus LAMs, organize their internal data when they are processing real-time visual game frames and text prompts.
Yuki: For the wider implications, this research suggests that we can use these complex AI systems as a tool to probe the specific neural circuits that handle planning and action selection in ways that might be inaccessible through traditional neuroimaging alone.
Ines: What this means in simpler terms is that we are getting a better way to understand the biological mechanisms of planning by seeing how well these sophisticated models mimic or differ from human brain activity during these complex tasks.
Marcus: It points toward a future where we can develop AI systems that aren't just good at predicting an outcome, but systems whose internal workings are structurally aligned with the actual computational pathways used by biological organisms for decision-making and behavior.
Yuki: That could provide insights into how cognitive structures have been shaped through evolutionary pressures, showing us which organizational patterns in AI mirror the adaptive strategies we see in living species.
Ines: It really moves the conversation from just observing brain activity to understanding the computational structures that generate that activity, whether it’s in a human or an AI system.
Marcus: So, ultimately, this paper contributes to understanding how predictive models function by linking them directly to specific cortical regions responsible for motor control and policy-relevant computations during active interaction with an environment.
Yuki: It’s a valuable piece of work because it shows the potential for computational modeling to inform our fundamental theories about neural computation and adaptive behavior in living systems.
Ines: And that’s where we leave it for now, connecting the findings on model structure to the broader implications for understanding cognitive biology and AI development.
Microsoft Research Bangalore, India · AWS AI Labs Amazon, USA
q-bio.NC, cs.AI, cs.LG
Submitted: 2026-05-19
Updated: 2026-10-07
Comments: 32 pages, 20 figures
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 83/100
The gist: Understanding how humans and artificial intelligence systems predict and plan by interacting with their environment is a fundamental challenge at the intersection of neuroscience and machine learning.
Key concepts
- Voxel-wise Encoding Performance
- This measures how well the brain activity patterns of an AI model match the actual neural activity in specific brain regions (voxels). The study found that both VLMs and LAMs showed significantly better voxel-wise encoding than standard reinforcement learning methods, suggesting they learn more action-relevant representations.
- Prompt-Driven Gains
- This examines how providing specific instructions (prompts) changes the AI's internal brain representation. The largest improvements occurred in higher-order areas like frontal and parietal regions, indicating that prompting recruits the same brain networks used for planning in traditional reinforcement learning.
- Variance Partitioning
- This technique looks at how different types of information (like action versus reasoning) are organized within the model's representation. The finding showed that VLMs are prompt-symmetric, while LAMs are action-asymmetric, meaning action fine-tuning captures more unique variance than reasoning prompts.
Terminology
Summary
Understanding how humans and artificial intelligence systems predict and plan by interacting with their environment is a fundamental challenge at the intersection of neuroscience and machine learning. The gist: action-specialized fine-tuning reorganizes multimodal representations toward action-relevant neural computations even when whole-brain prediction accuracy is statistically equivalent between VLM and LAM.
Model Comparison and Baseline Performance
The study investigates the brain alignment of two foundation model families—vision-language models (VLMs) and large-action models (LAMs)—during naturalistic Atari-style video games, comparing their performance against reinforcement learning (RL) baselines. Both VLMs and LAMs exhibit significantly exhibit voxel-wise encoding performance than RL baselines, with the advantage holding even under matched feature dimensionality.
Specifically, when comparing Qwen2.5-VL and InternVL3 against EMPA and DDQN baselines, both models significantly outperform the classical baselines in whole-brain voxel-wise encoding. Furthermore, even when VLM and LAM features are reduced to 8 dimensions to match EMPA’s feature count, they continue to significantly show higher brain alignment than both RL baselines,
ruling out feature count as the source of the gain. Performance saturates between 64-dim and 1024-dim, suggesting that the gain over RL baselines reflects what the features encode rather than feature dimensionality alone.
Prompting Conditions and Cortical Hierarchy
The researchers examine how action-focused and reasoning-focused prompts shape internal representations. The results show that prompt-driven gains scale with the cortical processing hierarchy: the largest improvements appear in frontal-parietal and motor-planning regions, while early visual cortex gains roughly half as much.
For instance, in ROI analysis, MFG shows the largest gain (∆ r = +0.189), followed by SMA (+0.182), IFGtriang and IFGoperc (both +0.149), and AG (+0.123).
This pattern suggests that prompt-driven alignment recruits the same higher-order regions that support theory-based RL.
However, the analysis of Qwen3.5 reveals a gap: generated reasoning-trace representations are substantially less brain-aligned than action-output representations,
indicating a substantial reasoning-action asymmetry in thinking-mode brain alignment.
Representational Organization via Variance Partitioning
A key finding is that while VLMs and LAMs achieve statistically equivalent whole-brain prediction,
variance partitioning reveals a qualitatively different representational organization. For VLMs, they are described as prompt-symmetric (12.5% unique action vs. 13.6% unique reasoning),
whereas LAMs are prompt-asymmetric (27% unique action vs. -5% unique reasoning), with the asymmetry strongest in frontal-motor cortex.
This dissociation is crucial: action fine-tuning already captures much of the information reasoning prompts provide,
as seen in the LAMs where action prompts contribute substantially more unique variance than reasoning prompts.
Cross-Family Generalization and Conclusion
To verify these findings, cross-family generalization was tested by comparing OSAtlas-Pro (LAM) with InternVL3 (VLM). The qualitative pattern is preserved: OS-Atlas shows strong action-asymmetry (uAction = 0.080, uReasoning = 0.017); InternVL3 is approximately prompt-symmetric (uAction = 0.060, uReasoning = 0.065).
This suggests that the within-family Qwen2-VL pattern reflects a general property of action-tuned vs general-purpose multimodal foundation models.
Ultimately, the study concludes that action fine-tuning reorganizes model representations toward motor-control and policy-relevant neural computations, despite similar overall predictive accuracy.
Summary of Key Findings
-
Both VLMs and LAMs significantly outperform RL baselines in voxel-wise brain encoding under no-prompt conditions.
-
Prompt gains are strongest in higher-order frontal-parietal and motor-planning regions (MFG, SMA, IFG).
-
Variance partitioning reveals that VLMs are prompt-symmetric, while LAMs are action-asymmetric and action-dominant in frontal/motor cortex.
-
Explicit reasoning traces from thinking models align less with brain activity than final action outputs.
Limitations
The study notes limitations, including the comparison being restricted to the Qwen2-VL family for matched architectures and the preliminary nature of testing reasoning traces on a single model (Qwen3.5). Furthermore, behavioral analyses like button-press latencies are not directly related to model representations in this work. The future direction involves systematic evaluation across multiple thinking-mode models with controlled trace-length and matched prompts.
Improvements for AI systems
As a fastidious researcher, I have thoroughly analyzed this paper, Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay.
The findings provide a critical bridge between cognitive neuroscience (fMRI) and modern artificial intelligence (VLMs/LAMs), specifically revealing how fine-tuning for action reorganizes internal representations compared to general-purpose reasoning.
Here are the specific, actionable improvements for AI systems derived from this research:
- [[Action-Specialized Fine-Tuning for Policy Reorganization]]
The paper shows that Large Action Models (LAMs) exhibit a qualitatively different representational organization than Vision-Language Models (VLMs). LAMs are prompt-asymmetric,
showing a strong preference for action features, especially in frontal-motor cortex.
Improvement: Implement an explicit fine-tuning regime where policy models are trained not just on reward signals, but specifically with prompts that emphasize action plans
and threat/goal consideration
(similar to the Action Prompt template). This targeted training should force the model to reorganize its internal representations toward action-relevant neural computations, even when raw prediction accuracy is equivalent to a VLM.
Improved AI Capability: LAMs will become superior at complex, goal-directed tasks requiring immediate motor planning and reactive decision-making (e.g., high-speed navigation in dynamic environments, real-time tactical game playing) because their internal structure will be pre-aligned with the motor control hierarchy observed in the brain.
- [[Hierarchical Prompting for Goal/Theory Alignment]]
The research demonstrates that prompt gains scale with cortical hierarchy: largest improvements occur in frontal-parietal and motor-planning regions, while early visual cortex gains are smaller. Furthermore, action prompts and reasoning prompts elicit distinct functional activations across the cortex (e.g., VLM favors reasoning; LAM favors action).
Improvement: Develop a Cognitive Steering
architecture for foundation models during interactive tasks. When the agent is in a complex decision-making phase (requiring high-level planning), it should be steered toward using reasoning prompts
to engage frontal-parietal networks. During immediate, reactive phases, it should switch to action prompts.
Improved AI Capability: This creates a more robust and flexible agent capable of switching between long-horizon strategic planning (reasoning) and rapid, high-fidelity execution (action), mimicking the human brain's ability to engage different cognitive networks on demand.
- [[Representation Decomposition for Interpretability]]
Variance partitioning revealed that while VLMs are prompt-symmetric, LAMs exhibit action-asymmetry where reasoning features become redundant or even detrimental in frontal-motor regions (negative unique reasoning variance).
Improvement: Integrate a Redundancy Detector
module into the model's internal representation layer. This module would continuously monitor the contribution of reasoning versus action prompts to the overall prediction error and flag situations where a specific prompt type is actively redundant or conflicting with an action-tuned structure.
Improved AI Capability: This provides unprecedented interpretability for complex agents, allowing developers to understand precisely when a model is relying on redundant information (e.g., when a high-level reasoning prompt adds no predictive value because the action policy already subsumes it), leading to more efficient and less computationally expensive models in production.
- [[Thinking-Mode Alignment Verification]]
The preliminary study on Qwen3.5 showed that the generated reasoning trace
representations align significantly less with brain activity than the final action output
representations, suggesting explicit chain-of-thought reasoning may not enhance brain alignment compared to direct action commitment.
Improvement: Prioritize training and evaluation of future thinking-mode models (like Qwen3.5) by focusing on aligning the final output/action commitment representation rather than the intermediate reasoning trace representation with fMRI data. If a model's final action output aligns well, it suggests that the thinking
process itself might be an internal, non-neural computation that does not need to be explicitly mapped to brain activity for behavioral alignment.
Improved AI Capability: This allows for the deployment of more efficient reasoning architectures where explicit, verbose reasoning traces are discarded in favor of compact, action-aligned decision vectors.
Abstract
Understanding how humans and artificial intelligence systems predict and plan by interacting with their environment is a fundamental challenge at the intersection of neuroscience and machine learning. Most brain-encoding studies focus on aligning artificial models with brain activity during language comprehension or passive visual processing, while interactive brain alignment studies have to date been largely limited to reinforcement-learning (RL) agents and theory-based models. To address this gap, we study brain alignment of representative models from two foundation-model types, namely vision-language models (VLMs) and large-action models (LAMs), using fMRI recordings from participants playing naturalistic Atari-style video games. Specifically, we examine how action-focused and reasoning-focused prompts shape the models' internal representations and their alignment with fMRI brain activity. First, we find that both VLMs and LAMs achieve significantly higher voxel-wise encoding performance than RL baselines, with the advantage holding even under matched feature dimensionality. Second, compared to a no-prompt baseline, prompt-driven gains are larger in higher-order frontal-parietal and motor-planning regions than in early visual cortex, roughly 2-2.5 times when averaged over region-of-interest (ROI) groups, although individual regions are heterogeneous. Third, variance partitioning reveals a qualitatively different representational organization. VLM representations are prompt-symmetric (12.4% unique action vs. 9.5% unique reasoning), whereas LAM representations are action-dominant (25.6% unique action vs.-8.2% unique reasoning), with the asymmetry strongest in frontal-motor cortex. Together, these results associate action specialization with distinct cortical alignment patterns in multimodal game-state representations, revealing differences hidden by similar prediction accuracy.
Sources
- Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners
- Vision-Language Integration in Multimodal Video Transformers (Partially) Aligns with the Brain
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- Human-Level Reinforcement Learning through Theory-Based Modeling, Exploration, and Planning
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- Atari-GPT: Benchmarking Multimodal Large Language Models as Low-Level Policies in Atari Games
- OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
Related papers
- BrainWave: A Brain Signal Foundation Model for Clinical Applications
- Toward Robust, Reproducible, and Widely Accessible Intracranial Speech Brain-Computer Interfaces: A Comprehensive Narrative Review of Neural Mechanisms, Hardware, Algorithms, Evaluation, Clinical Pathways and Future Directions
- CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution
- Emergence of psychopathological computations in large language models
- NeuroAI and Beyond: Bridging Between Advances in Neuroscience and Artificial Intelligence
- Attraction to hierarchical feature memory explains orientation bias