Brain-IT-VQA: From Brain Signals to Answers
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Brain-IT-VQA: From Brain Signals to Answers".
Jane: Decoding visual content from fMRI signals while answering questions about those images is a challenging problem, and this work introduces Brain-IT-VQA,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So we're diving into the paper Brain-IT-VQA: From Brain Signals to Answers today, which looks like it tackles the long-standing challenge of decoding what people are seeing from fMRI signals to answer questions about those images. Jane, you can give us the quick rundown on what this whole framework is trying to achieve?
Jane: Absolutely, Tom. Essentially, Brain-IT-VQA proposes a way to decode language tokens directly from brain activity and then use those tokens with a language model to answer visual questions. The core claim is that their method substantially outperforms previous fMRI-based captioning and VQA approaches. It’s about moving beyond just making captions; they want to use this data as a tool for actually understanding how different brain regions encode visual information (<ref:2605.29588#pg0>).
Lu: I find the idea of decoding language tokens from fMRI signals really intriguing, Jane. It suggests that we might be able to map specific linguistic concepts directly onto neural activity during visual perception. If we can get that mapping right, the possibilities for understanding visual processing architecture are massive (<ref:2605.29588#pg0>).
Meng: From an engineering standpoint, I'm more interested in how robust this end-to-end framework is when you integrate it with a pre-trained vision-language model like InstructBLIP, as the paper mentions. How does that integration actually handle the complexity of mapping brain activity to those language tokens?
Lalam: From my perspective as a language model, I see this as a fascinating way to ground visual understanding in biological reality. If we can reliably decode which brain activity corresponds to "object identity" versus "attributes," it could help refine how we structure knowledge representation within the AI itself (<ref:2605.29588#pg0>).
Tom: That’s a great point about grounding the understanding, Lalam. So, they aren't just guessing what the brain is doing; they are using a structured approach built on something called Brain-IT to predict those language tokens (<ref:2605.29588#pg0>). Jane, can you elaborate on what that initial prediction step involves without getting too deep into the technical math?
Jane: Certainly. The method starts by building upon the Brain Interaction Transformer, which they use to predict these language tokens from the fMRI signals, which they call BIT-L (<ref:2605.29588#pg0>). This step organizes the raw voxel-level signals into groups called "clusters of functionally similar voxels shared across subjects," which are then summarized into a "compact Brain Token."
Lu: Those compact Brain Tokens sound like a highly efficient way to distill massive amounts of fMRI data down to something meaningful before they even start interacting (<ref:2605.29588#pg0>). It’s an interesting compression strategy for neural information.
Paper summary: Meng: I'm thinking about the practical side of that tokenization, Lu. If those tokens are summarizing functional similarity, how does the model ensure it captures subtle visual details rather than just gross features? We need to make sure this works reliably on real-world data inputs (<ref:2605.29588#pg0>).
Lalam: The structure of these tokens could be really beneficial for cultural understanding, Meng. If we can isolate the brain activity related to "scene location" versus "object color," it helps us see how different aspects of visual experience get prioritized by the human brain (<ref:2605.29588#pg0>).
Tom: Speaking of structure, the paper introduces a two-pathway prediction system that’s quite clever. It uses a CLIP-aligned pathway and a direct conditioning pathway to generate prompt tokens for the language model (<ref:2605.29588#pg0>). Jane, how does this dual approach actually help them get better answers than just using one method?
Jane: That dual approach allows the model to leverage two different ways of relating brain activity to language. The CLIP-aligned pathway produces representations aligned with visual tokens, while the direct conditioning pathway predicts task-specific soft prompts for the language model directly from brain activity (<ref:2605.29588#pg0>). They then average these outputs to get the final prompt tokens needed for answering questions.
Lu: Averaging those two pathways seems like a smart way to balance pure visual alignment with direct task relevance, Jane. It suggests that both the visual structure and the immediate question context are important inputs when decoding what's happening in the brain (<ref:2605.29588#pg0>).
Meng: So, if we look at the results they present in this Brain-IT-VQA paper, Tom and Jane, what’s the most significant thing they claim about their performance compared to prior work? What’s the big win here?
Tom: The main point is that Brain-IT-VQA achieves state-of-the-art results on fMRI captioning and VQA directly from brain activity (<ref:2605.29588#pg0>). They show that their framework actually beats the previous methods they tested, which is a solid claim for any new technique in this area.
Jane: And what really sets them apart beyond just outperforming others is the introduction of NSD-VQA, which is a new dataset and benchmark specifically designed for fMRI analysis (<ref:2605.29588#pg2>). This allows researchers to do a controlled evaluation of question answering from brain activity in a way that was previously missing.
Lu: That NSD-VQA sounds like it opens up entirely new avenues for research, Jane. It moves the work from just showing performance gains to actively probing the organization of visual representations in the human brain (<ref:2605.29588#pg2>).
Meng: Probing how distinct types of visual and semantic information can be reliably inferred is what gets me interested from a practical application standpoint, Lu. It helps us understand which parts of the visual world are most easily accessible through brain signals (<ref:2605.29588#pg2>).
Paper summary: Lalam: And if we look at the analysis they did on decodable information, Tom and Jane, it’s quite revealing. They found a clear dependence on question type, where binary questions like "Y/N" achieve high accuracy around seventy-nine to ninety-three percent (<ref:2605.29588#pg2>).
Tom: That's a very concrete finding, Lalam. So, it suggests that the brain is much better at decoding coarse object presence or categorical distinctions when the question is simple and binary (<ref:2605.29588#pg2>). Jane, what about more complex questions?
Jane: For open-ended questions where you have to select among multiple semantic alternatives, like color or food, the performance drops quite a bit. They reported lower accuracy for categories such as color at forty-seven point eight three percent and action at sixty-six point three five percent (<ref:2605.29588#pg2>).
Lu: That drop in accuracy for fine-grained attributes like color is telling, Tom and Jane. It implies that those specific visual details require a much finer level of neural encoding than the broader concepts are (<ref:2605.29588#pg2>).
Meng: I’m thinking about the practical implication of that for things like image recognition systems we build with AI, Tom and Jane. If we only have good decoding for coarse categories, it tells us where our current visual models are strongest and where they still struggle to get precise details (<ref:2605.29588#pg2>).
Lalam: And the study also pointed out that scene-level questions, which ask about the overall context, remained highly accurate at ninety-three point zero zero percent (<ref:2605.29588#pg2>). This suggests that global contextual representations are more readily decoded than those fine-grained attributes we talked about earlier (<ref:2605.29588#pg2>).
Tom: That’s a key distinction, Jane, between scene context and specific attributes. And the analysis of voxel-cluster marginal contributions showed that different question categories engage distinct brain representations, like the "holding category" showing more concentrated activity in fewer regions (<ref:2605.29588#pg2>).
Jane: Exactly, Tom. That finding suggests that different types of visual and semantic information rely on partially different brain representations (<ref:2605.29588#pg0>). It’s not one single area doing everything; it’s specialized encoding (<ref:2605.29588#pg2>).
Lu: The masking-based analysis is really what makes this paper so valuable, Tom and Jane; it lets us actually probe the organization of visual representations in the human brain by seeing which regions contribute to which questions (<ref:2605.29588#pg2>). That opens up deep questions about how the brain organizes its visual knowledge structure.
Meng: If we take this forward, Lu, I wonder if this kind of functional mapping could inform how we design future multimodal AI systems—maybe creating models that are inherently structured around these biologically plausible representations instead of just statistical correlations (<ref:2605.29588#pg2>).
Lalam: For culture and creativity, Tom and Jane, this is huge because it gives us a biological blueprint for visual semantics. It helps us understand the fundamental ways humans process scenes before we even try to model them computationally (<ref:2605.29588#pg0>).
Paper summary: Tom: So, to wrap up on the implications of Brain-IT-VQA and NSD-VQA, Jane, what’s the big picture takeaway for us listeners about what this means for AI development in general?
Jane: The big picture is that we now have a framework that can be used to directly analyze how visual information gets processed in the brain when someone looks at an image (<ref:2605.29588#pg0>). It gives researchers a quantitative way to see which parts of the visual world are encoded by which parts of the brain, which is a step toward building more biologically informed AI systems (<ref:2605.29588#pg2>).
Lu: I think the future work mentioned in the paper, comparing these attribution results against known functional neuroimaging literature, will be crucial for validating these new findings and placing them correctly within the broader neuroscience context (<ref:2605.29588#pg2>). That validation is where we turn a strong finding into established knowledge.
Meng: From an engineering perspective, I’m excited about the potential to create more interpretable AI tools, Tom and Jane. If we can start mapping high-level visual concepts directly back to specific brain activity patterns, it gives us a new kind of debugging mechanism for complex vision models (<ref:2605.29588#pg2>).
Lalam: And for the cultural impact, this means that future AI systems won't just be pattern matchers; they could potentially build representations that mirror the structural organization of human visual experience, which is incredibly rich and complex (<ref:2605.29588#pg0>).
Tom: So, to summarize for our listeners: Brain-IT-VQA provides a method to decode language tokens from fMRI signals to answer visual questions, and NSD-VQA gives us the tools to systematically evaluate what kind of visual information is being decoded from brain activity (<ref:2605.29588#pg2>).
Jane: And the implication is that we can now probe which forms of visual and semantic information are reliably decoded from fMRI responses to natural images, allowing us to see how different functional brain regions support distinct types of questions (<ref:2605.29588#pg2>).
Lu: It’s a powerful tool for understanding the architecture of human visual cognition, Tom and Jane. That kind of detailed mapping is something we’ve only dreamed of being able to do this directly from brain signals (<ref:2605.29588#pg0>).
Meng: We need to keep an eye on how these findings translate into designing more nuanced AI architectures, because understanding that the brain prioritizes scene context over fine attributes is something we can actually use to build better models (<ref:2605.29588#pg2>).
Lalam: This work really moves the conversation from just "can an AI answer this?" to "how does a human brain process and structure this visual information?" which is a deeper level of inquiry for cultural and creative AI development (<ref:2605.29588#pg0>).
Conclusion: Tom: So, we've seen how Brain-IT-VQA decodes language tokens from brain activity to answer visual questions, and now we're wrapping up with the big picture on what that actually means.
Jane: That framework takes fMRI data and turns it into answers for pictures, which is a really neat way to look at how our brains process vision.
Lu: I think the title itself, Brain-IT-VQA, really captures the essence of taking something complex like brain signals and making them useful for answering questions.
Meng: From an engineering standpoint, the authors are using this to create a new benchmark called NSD-VQA to actually test how well we can decode those visual concepts reliably.
Lalam: I see the title as suggesting a direct bridge from biology—the brain—to what we can actually understand about images through language.
Tom: Exactly, and the authors are showing us that this isn't just some abstract concept; they’ve put it into practice with concrete results on how different visual information gets decoded.
Jane: They're suggesting that by using these brain signals, we can start to map out which specific areas of our brains handle which types of visual questions.
Lu: That mapping ability is huge because it helps us understand the underlying organization of human visual knowledge, and I'm really excited about how this could inform future AI architectures.
Meng: If we can actually see which brain regions are responsible for identifying an object versus its color, that gives us a new kind of structure to aim for in our own models.
Lalam: It means future AI systems won't just be pattern matchers; they could potentially build representations that mirror the structural organization of human visual experience, which is incredibly rich and complex.
Tom: It really puts things into perspective when you think about how we can move from just seeing what an AI can do to understanding the biological mechanisms behind it.
Jane: It's a powerful step in understanding the cognitive side of computer vision, showing us that there are tangible connections between neural activity and visual concepts.
Lu: We should really focus on those future work plans they mentioned, especially comparing these attribution results against established neuroimaging literature to truly ground this research in the wider field.
Meng: That validation is what will make this research move from a compelling study to something that actually influences how we build next-generation systems.
Lalam: It’s about making sure that the insights we gain about visual structure are robust and applicable across different modalities, which is vital for developing more versatile AI tools.
Weizmann Institute of Science
cs.CV, cs.AI, q-bio.NC
Submitted: 2026-05-28
Updated: 2026-10-02
Project page: https://mcosarinsky.github.io/brain-it-vqa
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 83/100
The gist: Decoding visual content from fMRI signals while answering questions about those images is a challenging problem, and this work introduces Brain-IT-VQA, a framework that decodes language tokens from
Key concepts
- Brain-IT (BIT-L)
- This component predicts language tokens directly from fMRI signals. It organizes raw brain activity into 'compact Brain Tokens' representing functionally similar voxels across subjects, which are then used to extract task-relevant representations through self-attention mechanisms.
- Brain Tokens
- These are summarized representations of voxel-level fMRI signals, grouped into clusters of functionally similar voxels shared across subjects. These tokens interact via self-attention and cross-attention to capture specific visual or semantic information relevant to the question asked.
- NSD-VQA
- This is a new dataset and benchmark designed for controlled evaluation of question answering from brain activity. It provides many question-answer pairs per image across 20 categories, allowing researchers to reliably test which types of visual understanding can be decoded from brain responses.
- CLIP-aligned pathway
- One prediction method where query tokens attend to Brain Tokens to generate representations that align with CLIP's visual tokens. This pathway helps map brain activity directly onto the visual concepts understood by a pre-trained vision model, aiding in answering visual questions.
Terminology
Summary
Decoding visual content from fMRI signals while answering questions about those images is a challenging problem, and this work introduces Brain-IT-VQA, a framework that decodes language tokens from brain activity to answer visual questions and provides a new benchmark for studying how different brain regions encode visual information.
The gist
Brain-IT-VQA provides a framework for visual question answering from fMRI by decoding language tokens from brain activity and integrating them with a language model, substantially outperforming previous fMRI-based captioning and VQA approaches.
How it works
-
The method builds on the Brain Interaction Transformer (Brain-IT [1]) to predict language tokens from fMRI signals, denoted as BIT-L.
-
BIT-L organizes voxel-level fMRI signals into
clusters of functionally similar voxels shared across subjects,
which are summarized into acompact Brain Token.
-
These Brain Tokens interact through self-attention layers, and a cross-attention mechanism with learnable query tokens extracts
task-relevant representations
from them. -
In the Brain-IT-VQA framework, this extension (BIT-L) is integrated with a pretrained vision-language model like InstructBLIP [25].
-
The model employs two complementary prediction pathways: a CLIP-aligned pathway where query tokens attend to Brain Tokens to produce representations aligned with CLIP visual tokens, and a direct conditioning pathway that predicts
task-specific soft prompts for the language model
directly from brain activity. -
The final prompt tokens are obtained by averaging the outputs of both pathways, which, together with the textual query as a text prefix, condition the frozen language model to generate captions or answers.
Key Contributions and Evaluation
The authors introduce two major contributions: Brain-IT-VQA and NSD-VQA.
-
Brain-IT-VQA is presented as an
end-to-end visual question answering from fMRI
framework that achievesstate-of-the art performance on fMRI-based captioning and visual question answering.
-
NSD-VQA is a new dataset and benchmark designed for
controlled evaluation of question answering from brain activity,
providing on average 20 question-answer pairs per image across 20 controlled categories, whichdisentangle multiple levels of visual understanding.
-
The framework allows for the analysis of
which forms of visual and semantic information can be reliably decoded from fMRI responses to natural images
using this benchmark.
Decoding Performance Analysis
The study systematically analyzes decodable information by organizing questions into controlled categories: object identity, attributes, pose, position, location, category.
-
The results show a
clear dependence on question type,
where binary (Y/N) questions consistently achieve high accuracy (typically 79–93%), reflecting the decoding ofcoarse object presence and categorical distinctions.
-
Open-ended questions requiring selection among multiple semantic alternatives are more challenging, showing lower performance for categories such as color (47.83%), food (54.02%), and action (66.35%).
-
Scene-level questions remain highly accurate at 93.00%, suggesting that
global contextual representations are more readily decoded than fine-grained attributes.
-
The analysis of
voxel-cluster marginal contributions
reveals that different question categories engage distinct brain representations; for instance, theholding category exhibits more spatially concentrated contributions in a small number of regions,
while food questions appearmore distributed across ventral visual regions.
Ablation and Interpretation
The researchers conducted an ablation study to evaluate component necessity. They found that the Q-Former and external data augmentation contribute meaningful improvements, while removing BIT-L alignment or end-to-end training causes substantial degradation.
Furthermore, comparing Brain-IT VQA against InstructBLIP applied to reconstructed images shows that decoding answers directly from brain activity is more effective than first reconstructing the image and then applying a VQA model.
The masking-based analysis demonstrates that different brain regions contribute selectively to different question types, suggesting that different types of visual and semantic information rely on partially different brain representations.
This framework serves as a tool for probing the organization of visual representations in the human brain.
Conclusion
The paper concludes by presenting Brain-IT-VQA as a SotA framework for visual Captioning & VQA directly from fMRI, introducing NSD-VQA to enable reliable and interpretable evaluation, and providing initial results on attributing information content to functional brain regions via a masking-based analysis. Future work includes comparing these attribution results against known functional neuroimaging literature.
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed the Brain-IT-VQA: From Brain Signals to Answers
paper. The core innovation lies in creating a novel framework that decodes visual information directly from fMRI signals into natural language outputs, coupled with a highly structured benchmark for neuroscientific analysis.
Here are the specific improvements I propose for AI systems, categorized by capability enhancement:
)1. Enhanced Visual-Semantic Reasoning via Neuro-Decoding (Brain-IT-VQA Core Capability)
The system can move beyond standard Vision-Language Models (VLMs) that rely on direct image input to perform Neuro-Informed VQA.
Specifically, the improved AI system will be capable of:
"Given an fMRI signal recorded while a subject views an image, the system can generate answers to complex questions about that image (e.g., 'What is the pose of the person's legs?') by directly decoding relevant brain activity into language tokens and integrating them with a large language model (LLM)."
This capability is superior because it bypasses the need for an explicit, often imperfect, image reconstruction step. It allows for probing how specific visual and semantic concepts (like pose, color, or location) are encoded in the brain's functional architecture.
)2. Fine-Grained Semantic Discrimination (NSD-VQA Enabled Capability)
The integration of the NSD-VQA benchmark allows for training and fine-tuning systems to specialize in decoding specific visual features rather than just general scene descriptions.
Specifically, the improved AI system can be trained to:
"Accurately distinguish between subtle semantic categories, such as identifying whether an animal is present (binary question) versus answering a nuanced open-ended question about its action or attribute (e.g., 'What is the animal doing?')."
This improvement stems from training on 20 controlled categories, allowing the model to learn distinct decoding pathways
for different types of visual information—for instance, one pathway optimized for object identity and another for spatial relations.
)3. Interpretable Brain Representation Mapping (Contribution Analysis Enabled Capability)
The ability to perform voxel-cluster marginal contribution analysis allows the system to become a tool for understanding brain organization, not just prediction.
Specifically, the improved AI system can be capable of:
"Providing a quantitative map showing which specific functional brain regions support different types of visual and semantic understanding. For instance, it can demonstrate that 'holding' questions rely on spatially concentrated regions related to human-object interaction, while 'food' questions rely on distributed ventral visual regions."
This capability transforms the AI from a black-box predictor into a scientific probe, enabling researchers to discover novel functional mappings between brain anatomy and cognitive functions.
)4. Robustness Against Data Scarcity (Data Augmentation Enabled Capability)
The training setup includes augmenting the dataset with un-fMRI images using an Image-to-fMRI encoder. This allows for more generalized models.
Specifically, the improved AI system can be capable of:
Maintaining high performance and generalizability when tested on visual scenes or questions that were not present in the original fMRI training data, by leveraging learned mappings from un-fMRI data.
This addresses the critical limitation of limited subject-specific fMRI data by allowing the model to learn a more robust, generalized representation of visual input from the brain activity.
)5. Direct Decoding Efficiency (End-to-End Fine-Tuning Enabled Capability)
The end-to-end fine-tuning stage, combined with techniques like LoRA on both the BIT and Q-Former modules, allows for highly efficient adaptation to specific tasks.
Specifically, the improved AI system can be capable of:
Rapidly adapting its decoding mechanism to new question types or specific semantic domains with minimal retraining effort (using LoRA), while maintaining state-of-the-art performance on both captioning and VQA tasks.
This makes the deployment of fMRI decoding models highly practical for rapid clinical or research prototyping.
Abstract
Decoding visual content from fMRI signals recorded while a person views images, and specifically answering questions about the seen images, is a long-standing challenge. While significant progress has been made in recent years in visual question answering (VQA) from fMRI, performance remains limited. Moreover, although recent models can make increasingly accurate predictions, they have rarely been used as tools for understanding the structure of visual representations in the brain. We present Brain-IT-VQA, a framework for visual question answering from fMRI. Unlike previous methods, which extract a fixed representation from the fMRI signal, our extraction is conditioned on the question itself, so what is decoded from brain activity depends on what is being asked. Our model substantially outperforms previous fMRI-based captioning and VQA approaches. We further introduce NSD-VQA, a new dataset and benchmark for visual question answering from fMRI. Unlike existing image-fMRI VQA datasets, which typically provide only a few broad and weakly controlled questions per image, NSD-VQA provides on average 20 question-answer pairs per image across 20 controlled question categories that disentangle multiple levels of visual understanding. This enables more reliable and interpretable evaluation despite limited fMRI test data. Together, Brain-IT-VQA and NSD-VQA provide both a strong predictive framework and a tool for studying brain representations. Using this benchmark, we quantify which forms of visual and semantic information can be reliably decoded from fMRI responses to natural images. We further analyze the contributions of different brain regions across question types.
Sources
- Brain Captioning: Decoding human brain activity into images and text
- UniBrain: A Unified Model for Cross-Subject Brain Decoding
- MindGPT: Interpreting What You See with Non-invasive Brain Recordings
- GPT-4 Technical Report
- Gemini: A Family of Highly Capable Multimodal Models
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- MindEye2: Shared-Subject Models Enable fMRI-To-Image With 1 Hour of Data
- The Wisdom of a Crowd of Brains: A Universal Brain Encoder
- BrainExplore: Large-Scale Discovery of Interpretable Visual Representations in the Human Brain
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Gemma: Open Models Based on Gemini Research and Technology
- The Llama 3 Herd of Models
- The Algonauts Project 2023 Challenge: How the Human Brain Makes Sense of Natural Scenes
- The Color of the Cat is Gray: 1 Million Full-Sentences Visual Question Answering (FSVQA)
- UniBrain: Unify Image Reconstruction and Captioning All in One Diffusion Model from Human Brain Activity
- Brain-language fusion enables interactive neural readout and in-silico experimentation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models