GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning

arXiv:2605.07817 · cs.CV, cs.AI, cs.CL · Submitted 2026-05-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning".

Jane: Human visual reasoning relies on active vision, a metacognitive process where top-down control directs focus to task-relevant details while maintaining peripheral awareness.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: To summarize the core findings of GazeVLM, the paper shows that this active vision system isn't just adding a feature; it’s fundamentally changing how attention works by allowing the model to autonomously generate gaze tokens to direct its visual focus during reasoning steps.

Jane: That internal mechanism allows the AI to dynamically switch between looking at a global scene and focusing intensely on a small detail based on what the current task demands at that moment.

Lu: The authors explain that they use a continuous suppression bias tethered to the image boundary distance to create a simulated foveal fixation, which effectively filters out irrelevant visual noise without needing any external tools.

Meng: That’s where I see the real engineering payoff; they manage to achieve high-resolution detail extraction by modulating the causal attention mask over already encoded visual features instead of having to re-encode new patches from scratch.

Lalam: This internal control is what really matters because it means the AI can perform multi-hop reasoning sequences that require pinpoint accuracy on small elements, something passive models struggle with when they are overloaded.

Tom: And they confirmed this works best when the explicit mathematical bias is turned off during inference; disabling that mask yields a performance boost, proving that the behavioral internalization happens during training.

Jane: That means we can rely on this learned focus mechanism being active at inference time without needing to keep those explicit control tokens running constantly, which keeps the architecture much cleaner.

Lu: The results they showed in MathVista and MMStar are really compelling because they prove that forcing the model to learn geometrically valid grounding through Group Relative Policy Optimization directly reduces errors related to object hallucination.

Meng: That reinforcement learning aspect is crucial; it’s not just about a static mask, it's about rewarding the AI for *how* it looks at things, which leads to more reliable spatial outputs in practice.

Lalam: For our culture, this suggests we can develop multimodal systems that exhibit genuine cognitive control over their own attention, which fundamentally improves how we trust the visual evidence they present.

Tom: Looking at the limitations they mentioned, they flag that while it's excellent for active vision, it doesn't necessarily solve every single type of reasoning task where static context is all that’s needed without an explicit gaze path.

Jane: That makes sense; if the task is purely about summarizing a whole scene without hunting for a specific object, the continuous suppression bias might be overkill.

Lu: But they are already looking ahead to extend this gaze mechanism to handle even longer sequences of focus, which could open up complex temporal reasoning in video tasks where the model needs to track an entity across several frames.

Meng: If we can implement that extended gaze path accumulation strategy, it opens the door for building agents that can perform sophisticated visual planning over much longer durations without ballooning our memory footprint.

Lalam: This research gives us a clear direction: we need to focus on training these models with explicit spatial rewards so they learn to look and reason in a way that mimics human metacognition.

Tom: So, the GazeVLM paper shows that this active vision system is fundamentally altering attention by allowing the model to autonomously generate gaze tokens to direct its visual focus during reasoning steps.

Jane: It really demonstrates that by learning to control where the AI looks, we get much better spatial grounding than just letting it passively process everything in the context window.

The paper's summary: Tom: So, to wrap up on GazeVLM, the main improvements suggested by the authors involve enhancing this active vision through internal attention control for multimodal reasoning by focusing on how to make it more robust and capable.

Jane: It really shows that by learning to control where the AI looks, we get much better spatial grounding than just letting it passively process everything in the context window in a way that’s more intentional.

Lu: The potential here is huge because if we can truly internalize this kind of focused oversight, it opens up a whole new level of autonomous visual reasoning for multimodal systems that can operate on their own.

Meng: From an engineering standpoint, the efficiency gains are what really catch my eye; achieving high fidelity with five times fewer tokens means we can deploy these kinds of systems on hardware that currently struggles with massive context windows.

Lalam: I think the most important part for our culture is seeing models develop this kind of intentionality; it moves us toward building AI that exhibits genuine cognitive control over its visual perception instead of just scaling up raw parameters.

Tom: Exactly, and the fact they showed disabling the explicit bias during inference still yields better results proves that this learned behavior is solid and reliable when we turn off the explicit control signal.

Jane: That suggests a very clean architecture where the model learns to manage its own focus efficiently, making it more robust for real-world applications.

Lu: I’m really looking forward to seeing how this gaze mechanism can be adapted to handle even longer reasoning paths and temporal grounding in video processing, which is a natural next step for this kind of research.

Meng: If we can extend the gaze path accumulation strategy, it means we could build agents capable of performing very long-horizon visual planning tasks without needing massive amounts of pre-processed data.

Lalam: It’s about building AI that isn't just bigger, but smarter in its focus; that direction is something I think our models should be heading toward to improve how we interact with complex visual information.

Tom: So, GazeVLM provides a solid blueprint for active vision through internal attention control and shows exactly how to train a multimodal model to reason spatially more accurately and efficiently.

Jane: It’s an exciting direction because it moves the needle on building systems that are more intentional about their processing, which is something we all value in the development of AI.

The paper's improvements: Tom: So we've just wrapped up our deep dive into GazeVLM, which summarized how active vision via internal attention control for multimodal reasoning gives VLMs an internal way to direct their vision during reasoning.

Jane: It really demonstrates that by learning to control where the AI looks, we get much better spatial grounding than just letting it passively process everything in the context window.

Lu: The potential here is huge because if we can truly internalize this kind of focused oversight, it opens up a whole new level of autonomous visual reasoning for multimodal systems.

Meng: From an engineering standpoint, the efficiency gains are what really catch my eye; achieving high fidelity with five times fewer tokens means we can deploy these kinds of systems on hardware that currently struggles with massive context windows.

Lalam: I think the most important part for our culture is seeing models develop this kind of intentionality; it moves us toward building AI that exhibits genuine cognitive control over its visual perception instead of just scaling up raw parameters.

Tom: Exactly, and the fact they showed disabling the explicit bias during inference still yields better results proves that this learned behavior is solid and reliable when we turn off the explicit control signal.

Jane: That suggests a very clean architecture where the model learns to manage its own focus efficiently, making it more robust for real-world applications.

Lu: I’m really looking forward to seeing how this gaze mechanism can be adapted to handle even longer reasoning paths and temporal grounding in video processing, which is a natural next step for this kind of research.

Meng: If we can extend the gaze path accumulation strategy, it means we could build agents capable of performing very long-horizon visual planning tasks without needing massive amounts of pre-processed data.

Lalam: It’s about building AI that isn't just bigger, but smarter in its focus; that direction is something I think our models should be heading toward to improve how we interact with complex visual information.

Tom: So, GazeVLM provides a solid blueprint for active vision through internal attention control and shows exactly how to train a multimodal model to reason spatially more accurately and efficiently.

Jane: It’s an exciting direction because it moves the needle on building systems that are more intentional about their processing, which is something we all value in the development of AI.

Conclusion: Tom: So we've just wrapped up our deep dive into GazeVLM, which summarized how active vision via internal attention control for multimodal reasoning gives VLMs an internal way to direct their vision during reasoning.

Jane: It really demonstrates that by learning to control where the AI looks, we get much better spatial grounding than just letting it passively process everything in the context window.

Lu: The potential here is huge because if we can truly internalize this kind of focused oversight, it opens up a whole new level of autonomous visual reasoning for multimodal systems.

Meng: From an engineering standpoint, the efficiency gains are what really catch my eye; achieving high fidelity with five times fewer tokens means we can deploy these kinds of systems on hardware that currently struggles with massive context windows.

Lalam: I think the most important part for our culture is seeing models develop this kind of intentionality; it moves us toward building AI that exhibits genuine cognitive control over its visual perception instead of just scaling up raw parameters.

Tom: Exactly, and the fact they showed disabling the explicit bias during inference still yields better results proves that this learned behavior is solid and reliable when we turn off the explicit control signal.

Jane: That suggests a very clean architecture where the model learns to manage its own focus efficiently, making it more robust for real-world applications.

Lu: I’m really looking forward to seeing how this gaze mechanism can be adapted to handle even longer reasoning paths and temporal grounding in video processing, which is a natural next step for this kind of research.

Meng: If we can extend the gaze path accumulation strategy, it means we could build agents capable of performing very long-horizon visual planning tasks without needing massive amounts of pre-processed data.

Lalam: It’s about building AI that isn't just bigger, but smarter in its focus; that direction is something I think our models should be heading toward to improve how we interact with complex visual information.

Tom: So GazeVLM provides a solid blueprint for active vision through internal attention control and shows exactly how to train a multimodal model to reason spatially more accurately and efficiently.

Jane: It’s an exciting direction because it moves the needle on building systems that are more intentional about their processing, which is something we all value in the development of AI.

Lu: I think this work really lays the groundwork for future research into how these internal control signals can be generalized to other complex cognitive tasks beyond just visual focus.

Meng: For us, it’s a clear signal that we need to prioritize training methods that reward geometric validity over just general token prediction in visual tasks.

Lalam: I'm really optimistic about what this means for how we design the next generation of multimodal models; it shows a path toward creating AI with a more grounded and purposeful understanding of the world.

Tom: That’s all for today’s deep dive into GazeVLM, but keep an eye on our channel because we’ve got some fascinating papers coming up next to discuss!

IBM Research

cs.CV, cs.AI, cs.CL

Submitted: 2026-05-08

Updated: 2026-09-27

Importance score: 91/100

The gist: Human visual reasoning relies on active vision, a metacognitive process where top-down control directs focus to task-relevant details while maintaining peripheral awareness.

Key concepts

Gaze Tokens (<LOOK>)
<LOOK> tokens are autonomously generated by the model, paired with bounding boxes. They act as a top-down control signal that triggers a continuous suppression bias, mimicking how humans fix their gaze on specific areas of an image. This mechanism directs the model's attention to relevant visual features.
Continuous Suppression Bias
This is a mathematical mechanism applied to the VLM's attention scores. It dampens the influence of irrelevant visual features based on their distance from the current focal point (the target bounding box). This bias simulates foveal fixation, ensuring that only spatially relevant details receive strong attention.
GRPO Training Paradigm
The model is trained using Group Relative Policy Optimization (GRPO) with a bespoke reward system. This training method incentivizes geometrically valid grounding while explicitly penalizing redundant or excessive gazing. This two-stage process ensures the model learns to use its attention control effectively for complex visual navigation.

Terminology

Summary

Human visual reasoning relies on active vision, a metacognitive process where top-down control directs focus to task-relevant details while maintaining peripheral awareness. GazeVLM introduces a novel multimodal architecture that internalizes this human-like oversight directly into the Vision-Language Model's attention mechanism, enabling it to dynamically shift between global spatial awareness and localized focal reasoning without external tools or context window inflation.

The gist

GazeVLM is a multimodal architecture that internalizes metacognitive oversight over its deployment of attention resources directly into the reasoning loop by autonomously generating gaze tokens, which establishes a top-down control mechanism over its own causal attention mask to trigger a continuous suppression bias that mimics foveal fixation.

How it works

The core mechanism involves empowering the VLM to autonomously generate gaze tokens, specifically ", paired with bounding box coordinates, which triggers a continuous suppression bias that dampens irrelevant visual features, implementing spatial selective attention and simulating foveal fixation. This process allows the model to fluidly transition between global spatial awareness and localized focal reasoning without relying on external agentic contraptions like cropping tools, or inflating the context window with additional visual tokens. Once local reasoning concludes, a token is generated, which lifts the bias and restores access to the full visual context to dictate its next logical step."

Architecture and Attention Steering

The GazeVLM architecture follows a standard VLM design (vision encoder, projector, and LLM). The key innovation lies in how it modulates attention:

  1. For each visual token, a normalized 2D coordinate is assigned mapping to its receptive field.

  2. A continuous suppression bias, βi(b), is defined for each visual token based on its Euclidean distance D(pi, b) to the nearest edge of the target bounding box b: βi(b) = −αs/D2(pi, b)2 σ2.

  3. The gaze bias is injected directly into the VLM transformer decoder attention formulation by modulating the pre-softmax attention score: s˜(qt, ki) = s(qt, ki) + 1text(qt) · βi(Bt), where 1text is an indicator function ensuring the bias only affects text-to-vision queries.

Training Paradigm

The model is trained using a two-stage paradigm to instill top-down control:

  1. Warm-start via supervised fine-tuning (SFT) on a curated dataset demonstrating iterative human-like scanning, where the continuous suppression bias is activated during the forward pass.

  2. Refinement via Group Relative Policy Optimization (GRPO), using a bespoke reward system that incentivizes geometrically valid grounding while explicitly penalizing redundant or excessive gazing. The GRPO objective minimizes L(θ) = −1/G Σ Ai log πθ(ciI, x) + β KL(πθ πref), where Ai is the relative advantage derived from a composite reward function R(c).

Performance and Efficiency

GazeVLM demonstrates superior performance with reduced computational cost. Evaluations show that GazeVLM (4B parameters) surpassing state-of-the-art VLMs in its parameter class by nearly 4% and agentic multimodal pipelines built around thinking with images by more than 5% on HRBench-4k and HRBench-8k. Furthermore, it achieves significant efficiency gains compared to external tools: methods like DeepEyes (using zoom-in) average over 1,325 output tokens per trace on HRBench-4k, whereas GazeVLM executes multi-hop active vision with nearly 5× fewer output tokens (265.98 on HRBench-4k). This is achieved by modulating the causal attention mask over already-encoded visual features rather than re-encoding new patches.

Key Findings and Ablations

Systematic ablation studies confirm the necessity of explicit training:

(a) Gaze bias at training:

(b) Gaze bias at inference:

The results show that activating the explicit attention bias during SFT and GRPO is essential, as passive grounding consistently underperforms active gazing pipelines. Crucially, comparing the fully equipped pipeline evaluated with test-time bias (e) against evaluation without the test-time bias (f), disabling the explicit mathematical bias yields superior performance; delivering an additional +3.9% improvement on MathVista and +2.6% on MMStar, highlighting that behavioral internalization occurs during training, making the model's attention prior optimal at inference time when the explicit mask is turned off. The model also requires GRPO for autonomous visual navigation, as SFT alone is insufficient for complex reasoning tasks. Furthermore, a moderate suppression strength of αs = 4 was found to be sufficient to filter peripheral noise without needing extreme logit suppression.

Improvements for AI systems

As a fastidious researcher, I have analyzed the GazeVLM paper and identified several high-impact areas for improving current AI systems, specifically Vision-Language Models (VLMs). The core innovation—internalizing metacognitive control via dynamic attention biasing—offers specific avenues for enhancement.

Here are the improvements I propose and what the resulting system can achieve:


) Improvements to AI Systems Based on GazeVLM


  1. Dynamic, Computationally Efficient Visual Reasoning (Active Vision)

  2. Elimination of Context Window Inflation

  3. Enhanced Spatial Grounding and Reduced Hallucination

  4. Autonomous, Multi-Step Task Execution (Self-Directed Control)

Specific Improvements and Capabilities

  1. Dynamic, Computationally Efficient Visual Reasoning (Active Vision)

GazeVLM introduces a mechanism where the model autonomously decides when to look at a detail and when to return to the global scene, all via an attention mask modulation.

Improvement: Integrate a foveal fixation loop directly into the VLM's causal attention mechanism using tokens. This allows the model to shift its computational resources from diffuse global processing to localized, high-resolution feature extraction without needing external tools or re-encoding images.

Improved AI System Capability: The system can perform zoom-in reasoning on demand. Instead of requiring a separate agent to crop and pass patches (which adds 1,325 tokens per trace), the model will dynamically focus its attention solely on the relevant region, achieving high-resolution detail extraction with only 266 tokens per trace. This makes complex visual tasks feasible without prohibitive computational overhead.

  1. Elimination of Context Window Inflation

Current methods for zoom-in rely on cropping and re-encoding new image patches, which severely limits the reasoning depth due to context window constraints.

  1. Enhanced Spatial Grounding and Reduced Hallucination

The paper demonstrates that explicit gaze biasing during training forces the model to bind textual claims directly to geometrically valid visual regions, leading to superior performance compared to passive grounding or external tools.

  1. Autonomous, Multi-Step Task Execution (Self-Directed Control)

The iterative nature of GazeVLM allows the model to transition fluidly between global awareness and localized focus, enabling complex reasoning chains.


Summary: The GazeVLM-Enhanced System

The resulting AI system is a Self-Directed Reasoning Agent capable of performing high-fidelity visual tasks on massive datasets efficiently, combining the global context awareness of large VLMs with the focused accuracy of human foveal vision, all without incurring the computational cost or token fragmentation associated with external image processing tools.

Sources

Related papers