MG-VQA: Manipulation Grounded Visual Question Answering with VLMs

arXiv:2608.17129 · cs.CV, cs.RO · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "MG-VQA: Manipulation Grounded Visual Question Answering with VLMs".

Jane: Vision-language models (VLMs) are being pushed beyond static image analysis to handle real-world tasks that require physical interaction, such as answering questions about objects hidden in cluttered environments.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So we're diving into the paper "MG-VQA: Manipulation Grounded Visual Question Answering with VLMs," which really tackles how vision-language models handle real-world stuff like finding things that are actually hidden in a messy space. The core idea is formalizing this challenge as Manipulation Grounded Visual Question Answering, or MG-VQA, and introducing the PROBE framework to test and fine-tune agents for these dynamic tasks.

Jane: That sounds really practical, Tom. Basically, the paper argues that standard visual question answering isn't enough when you need an AI to actually interact with a physical environment to get the right answer. It highlights that answering questions about objects requires reasoning in situations where you have to move things around first to see what’s hidden behind them.

Lu: I think it's exciting because it pushes VLMs past just looking at a picture and forces them into planning sequences of actions, which opens up possibilities for much more complex embodied AI systems. The paper sets up a clear structure for evaluating this kind of interaction-based reasoning in cluttered scenes.

Meng: From an engineering standpoint, I'm interested in how they structured the evaluation to make sure these agents are actually learning useful manipulation skills rather than just getting lucky with perception tools.

Lalam: I think this research is significant because it shows how we can distill successful manipulation strategies from a much larger model into smaller, open-weight models that can generalize to new objects without needing retraining on every single type of thing.

Tom: Exactly, Lalam. The paper claims that agentic tool-based methods significantly outperform perception-only baselines across all task types in MG-VQA, with an average improvement of eight point zero percent <ref:2608.17129#pg1>. It's not just about seeing things; it's about using tools to change the scene and figure out the answer.

Jane: So, what makes this benchmark robust enough to give us a reliable measure of how good these agents are at manipulating objects in cluttered settings? The paper introduces PROBE as a comprehensive framework for that purpose.

Lu: PROBE has three main components: PROBE-Sim, which is this high-fidelity tabletop simulator with a UR5 robot and one thousand seven hundred ninety-six real-scanned everyday objects, along with manipulation primitives like grasp and push. It models physical dynamics where actions have non-zero failure probabilities.

Meng: A simulator with real objects is helpful for testing the practical aspects of deployment. But how do they ensure that the evaluation tasks in PROBE-Bench are truly representative of real-world clutter and reasoning needs?

Paper summary: Lalam: The PROBE-Bench suite has one hundred fifty tasks covering six question types, like Count, Find, Compare, Size, Reduce, and Beneath <ref:2608.17129#pg0>. These questions are algorithmically generated to test reasoning where the VLM might perceive objects before actually manipulating them.

Tom: It sounds like they designed a way to test if an AI can do the perception-grasp-push sequence before answering a question. That's a big step beyond just looking at a static image, isn't it?

Jane: It is, and the paper also presents PROBE-Agent, which is this fine-tuning recipe designed to take knowledge from a powerful teacher model like Gemini three point one Pro and distill it into smaller open-weight student models like Qwen3-VL or NVILA <ref:2608.17129#pg1>.

Lu: That distillation process seems key for making these complex skills accessible to smaller systems, especially open-weight ones, which is a really important direction for the AI community right now.

Meng: I see the focus on efficiency there. The recipe encourages manipulation-efficient question answering, meaning the model learns *when* and *where* to manipulate rather than just blindly invoking perception tools every time.

Lalam: And we see positive signs in those fine-tuned models; they match or approach the teacher's average accuracy, even showing positive gains on specific tasks like the Beneath task on held-out data, which suggests compositional generalization to unseen tasks.

Tom: So, to wrap up that part of the paper, they show that this agentic approach is effective and that a proper fine-tuning recipe can transfer those skills to smaller models for unseen objects. But what does this mean for the broader application of vision-language models?

Jane: It means we can move VLMs out of just analyzing single static images and into scenarios where they need to plan physical interactions to solve complex, real-world queries. It’s about giving them a physical agency that allows them to interact with their surroundings dynamically.

Lu: The potential for this is huge because it moves the capability from passive understanding to active problem-solving in embodied settings. Think about how much more useful an AI could be if it could reliably answer "Is my medication still in the cabinet?" by actually checking behind the containers.

Meng: Practically, I'm thinking about deployment constraints. The paper validates their sim-to-real transfer by deploying these finetuned policies on a Kinova Jaco robot in real-world settings, and Qwen3-VL-8B-SFT showed performance more than double its base model on that hardware without seeing real data during fine-tuning <ref:2608.17129#pg1>.

Lalam: That sim to real validation is crucial because it proves the learned policies aren't just simulator artifacts; they work when deployed on actual physical hardware, and the fact that it outperformed direct mode in the real world, where Gemini three point one Pro struggled with occluding distractors, is a strong indicator <ref:2608.17129#pg1>.

Paper summary: Tom: It really puts the power of agentic planning into perspective when you see those kinds of performance gains on real robots. It confirms that giving an AI tools to interact with the world can unlock capabilities that were previously out of reach for perception-only models.

Jane: The paper also points out some limitations, which is important context. They mention that the current toolshed in PROBE-Sim is limited to a discrete fixed tool library, which means the agents might struggle if they need to recover from failures by trying different recovery methods beyond what’s explicitly programmed.

Lu: That limitation on recovery from failures is something we definitely need to address in future work. A truly robust agent needs better tools for handling unexpected outcomes during manipulation.

Meng: From a practical deployment view, that fixed tool library is a constraint because real-world environments are rarely perfectly controlled or object-sparse; the agent needs more flexibility to handle novel objects and unforeseen obstacles.

Lalam: I think the research shows a strong path forward, especially with the fine-tuning recipes that show positive gains on unseen tasks like Beneath, suggesting that compositional generalization is achievable even in this manipulation domain.

Tom: So, while they have these limitations with the toolshed and recovery from failure, the overall message of MG-VQA is clearly about unlocking a new level of reasoning for VLMs through physical interaction and structured benchmarking.

Jane: The implication for the field seems to be that we need to shift our thinking toward agents that don't just see, but actively plan how to use tools and interact with their environment to find answers. It’s about modeling dynamic scenes where every action matters for the next step.

Lu: I think this opens up avenues for much more nuanced understanding of physical relationships within visual data, moving beyond simple object recognition to true scene comprehension.

Meng: For engineers building these systems, it means we need to think about designing robust tool interfaces and training regimens that encourage efficient planning rather than just brute-force perception loops.

Lalam: And for culture, I see this as a step toward developing AI that can handle complex physical tasks with more reliability and less reliance on perfectly curated training data for every single scenario.

Tom: That's a solid summary of where the paper lands, connecting the high-level goal of MG-VQA to the tangible benefits of tool use and fine-tuning for smaller models. We’ve covered a lot about what this research actually accomplishes in terms of testing and distillation.

Conclusion: Tom: So, to wrap up what we’ve seen today, we’re talking about MG-VQA: Manipulation Grounded Visual Question Answering with VLMs, and it really boils down to giving vision models the ability to actually interact with things in a messy environment.

Jane: It is a big step because it moves those models past just looking at pictures and into planning actions like grasping or pushing objects before they can answer a question about what’s hidden.

Lu: I think the real power here lies in how the framework PROBE sets up these rigorous tests, creating a simulation where the AI has to learn physical intuition instead of just memorizing visual patterns.

Meng: From an engineering standpoint, it’s fascinating that they're using a high-fidelity simulator with real objects to train these agents; that kind of grounding is essential for making them reliable when we actually put them into a workspace.

Lalam: I find the focus on distillation really compelling because it shows how we can take those complex physical reasoning skills from a massive foundation model and shrink them down into smaller, more accessible models that can be used widely.

Tom: Exactly! And the authors of this paper have laid out a pretty clear path showing how agentic tools lead to better results on these tasks compared to just relying on perception alone.

Jane: That means for the future, we aren't just training models to recognize things; we are training them to solve problems by physically manipulating those things.

Lu: The implication here is huge because it suggests a path toward AI that can truly reason about physical relationships in the real world, not just in digital representations.

Meng: Practically, this means we need to design better interfaces for these agents so they can efficiently use their tools without getting stuck on inefficient planning loops.

Lalam: I think the cultural impact will be significant because when AI can reliably handle complex physical tasks in novel settings, it opens up applications in everything from domestic assistance to advanced scientific discovery.

Tom: Speaking of new directions, we should also look at how this moves us toward training models that are more capable of generalizing those manipulation strategies to entirely new objects they haven't seen before.

Jane: That generalization capability is what makes this research so compelling for the broader field of machine learning development.

NVIDIA

cs.CV, cs.RO

Submitted: 2026-08-17

Updated: 2026-10-02

Project page: https://esi-bench.github.io

Importance score: 89/100

The gist: Vision-language models (VLMs) are being pushed beyond static image analysis to handle real-world tasks that require physical interaction, such as answering questions about objects hidden in cluttered

Key concepts

Manipulation-Grounded Visual Question Answering (MG-VQA)
This is a challenge where a model must answer questions about objects hidden in cluttered scenes by physically interacting with them. Instead of just looking at an image, the model needs to 'see,' 'grasp,' and 'push' objects to find the correct answer. It tests if VLMs can perform real-world tasks.
PROBE-Sim
This is a high-fidelity virtual laboratory featuring a UR5 robot and thousands of real objects. It allows VLMs to practice physical actions like grasping or pushing, simulating real physics and non-deterministic outcomes. This provides a safe environment for training agents on manipulation skills.
PROBE-Agent
This is a specific fine-tuning recipe designed to teach smaller open models how to use tools effectively. It combines direct observation with agentic rollouts, encouraging the model to plan deliberate sequences of actions rather than just calling perception tools repeatedly. This aims for efficient, successful manipulation planning.

Terminology

Summary

Vision-language models (VLMs) are being pushed beyond static image analysis to handle real-world tasks that require physical interaction, such as answering questions about objects hidden in cluttered environments. This research formalizes this challenge as Manipulation-Grounded Visual Question Answering (MG-VQA) and introduces PROBE, a comprehensive framework for benchmarking and finetuning VLM agents on these dynamic tasks. The work demonstrates that agentic tool-based methods significantly outperform perception-only baselines, and a fine-tuning recipe can distill successful manipulation strategies into smaller open-weight models with positive transfer to unseen objects.

The gist

Agentic tool-based methods outperform their perception-only baselines (8.0% on average) across all task types in Manipulation Grounded Visual Question Answering (MG-VQA).

PROBE Framework Components

The framework is built upon three main components designed to address the limitations of static VQA evaluation:

  1. PROBE-Sim: This is a high-fidelity tabletop simulator equipped with a UR5 robot and a toolshed that VLMs can invoke. It comprises 1,796 real-scanned everyday objects across 481 categories and provides manipulation primitives like grasp and remove and push. The simulator models physical dynamics, including non-deterministic outcomes for actions, where Each grasp and push has non-zero failure probability.

  2. PROBE-Bench: This is an evaluation suite of 150 tasks across six question types: Count, Find, Compare, Size, Reduce, and Beneath. These tasks are designed to test reasoning capabilities in cluttered scenes where a VLM may perceive, grasp, and push objects before answering. Questions are generated algorithmically using diverse prompt templates and simulator-derived ground-truth answers. Human verification ensures that all tasks are unambiguous and solvable, establishing an upper bound of performance at 100%.

  3. PROBE-Agent: This is a finetuning recipe designed to distill successful trajectories from a powerful teacher foundation model (Gemini 3.1 Pro) into a smaller open-weight student model (Qwen3-VL or NVILA). The training uses a mixed data recipe combining direct and agentic rollouts, encouraging manipulation-efficient question answering.

Key Findings on Tool Use and Generalization

The research investigates the impact of tool access and the efficacy of fine-tuning for open-weight models:

(1) Tool Access Improvement:

Across every frontier VLM we evaluate, granting tool access improves accuracy. Specifically, agentic VQA outperforms direct VQA by 8.0% on average (R1). This effect is most significant for weaker VLMs, as Opus 4.7 improves from 48.7% to 63.1% and Gemma-4 from 46.4% to 62.4%. However, direct VQA can achieve reasonable success on some tasks by exploiting partial visibility, though this is unreliable for manipulation-heavy tasks like Count and Reduce where full scene de-cluttering is needed before answering.

(2) When Tools Help:

The effectiveness of tools depends on the question type. Tools help when the question signals a clear manipulation target, but hurt when the agent must figure this out itself (R2). For tasks like Count and Reduce, Tools+wrong exceed Tools+correct for both models, indicating that incorrect target selection compounds across steps.

(3) PROBE-Agent Success:

The supervised fine-tuning recipe shows promise for smaller models. The resulting open-weight students match or approach the teacher's average accuracy, with positive gains on the held-out Beneath task (7.6% and 15.4% respectively), demonstrating positive signs of compositional generalization to an unseen task. Qualitative examples show that SFT teaches efficient tool-use planning, not just tool invocation, where the model plans a deliberate decluttering sequence rather than looping over perception tools and failing to grasp the target.

Sim-to-Real Transfer Validation

The study validates the learned policies by deploying PROBE-Agent finetuned models in real-world environments. Testing on a Kinova Jaco robot using a front-facing camera, Qwen3-VL-8B-SFT more than doubles its base model’s performance and approaches the teacher despite never seeing real-world data during fine-tuning. While the agentic mode outperforms direct mode in the real world—where Gemini 3.1 Pro struggles with occluding distractors—a performance gap remains between the finetuned model and the teacher, suggesting there is also potential in real world finetuning data to further close this gap.

Limitations

The current toolshed is limited to a discrete fixed tool library, which restricts recovery from failures.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on this research, categorized by capability:


  1. A significant improvement in Real-World Grounded Reasoning for Visual Question Answering (VQA) in cluttered environments.

  2. The ability for VLM agents to perform complex, multi-step physical manipulation tasks (e.g., picking up an object, moving it aside to reveal a hidden one) based on a single question, rather than requiring pre-labeled static images.

  3. Enhanced Compositional Generalization in smaller, open-weight models by distilling successful tool-use strategies from large teacher models (like Gemini 3.1 Pro) into efficient student policies (e.g., Qwen3-VL or NVILA). This makes high performance achievable even with less powerful foundational models.

  4. Improved Tool Use Efficiency by teaching agents to distinguish between situations where tools help (clear manipulation targets) and when they hurt (when the agent must figure out the necessary sequence itself), leading to better decision-making under uncertainty.

  5. Robust Sim-to-Real Transfer capability, allowing policies fine-tuned entirely in a high-fidelity simulation environment to perform effectively on real robots using front-facing cameras, bridging the gap between synthetic training and physical deployment.

The improved AI system can now:

  1. Answer complex, dynamic questions about physical states in real-world cluttered spaces (e.g., Is my medication still in the cabinet? or What is under the red bowl?).

  2. Execute a sequence of perception and manipulation actions autonomously to gain necessary visual information before formulating a correct answer (e.g., identifying, grasping, pushing objects).

  3. Achieve competitive performance on challenging VQA tasks using smaller, more accessible open-weight models by leveraging knowledge distilled from much larger teacher models.

  4. Adapt its tool-use strategy dynamically based on the question asked—using tools only when they provide a clear path to the answer rather than attempting complex, unnecessary manipulation sequences.

  5. Operate reliably on physical robots in real environments, demonstrating learned skills (from simulation) that generalize to novel objects and unseen scenarios.

Sources

Related papers