Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal LLMs
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal LLMs".
Tom: MLLMs currently attend to all visual tokens during generation, leading to diluted focus and unnecessary computational overhead, whereas human visual perception is inherently selective.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, to kick things off, let’s talk about the title and who put this paper together. The paper is called "Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal LLMs," and it features researchers like Junha Song, Byeongho Heo, Geonmo Gu, Jaegul Choo, Dongyoon Han, and Sangdoo Yun.
Jane: Those authors are definitely hitting on the core concept of selective attention; they’re tackling the problem of MLLMs attending to all visual tokens when human perception is inherently selective.
Lu: I think it's interesting how they frame it as query-adaptive routing rather than just a static selection process, which suggests a level of dynamic adaptation during the generation step itself.
Meng: Dynamic adaptation sounds promising for practical applications because in real-time scenarios, you can’t afford to waste computation on irrelevant visual data.
Lalam: I see the authors are trying to solve that diluted focus issue we've discussed; they want to make sure my attention is sharp and targeted, not just spread thin across the whole visual input.
Tom: Exactly, and when we look at the overall implication of this paper, it’s about moving past blanket attention toward a more human-like way of processing multimodal information during text generation.
Jane: It suggests that efficiency isn't just about cutting down on tokens; it’s about making the computation smarter by matching the model's focus to what is actually needed for the output.
The paper's summary: Tom: Now, let’s break down what Gaze Attention actually does according to the summary of "Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal LLMs." It explains that they spatially group visual embeddings into compact gaze regions, each summarized by a lightweight descriptor.
Jane: That grouping into gaze regions is key; it means the model doesn't look at individual pixels everywhere, but rather at these pre-summarized visual chunks to decide where to concentrate its attention.
Lu: The paper states that at every decoding step, the model dynamically selects the most relevant regions based on those descriptors and restricts attention only to them, which is what enables that focused attention on task-relevant regions.
Meng: That dynamic selection process sounds like it directly tackles the computational overhead problem we’ve been worried about; by restricting attention to a small subset of regions, it should significantly cut down the FLOPs involved in processing every visual token.
Lalam: And they don't just stop there, they add learnable context tokens appended to each image or frame so that even when most visual regions are hidden, these tokens provide a persistent global summary processed within the LLM’s native attention layers.
Tom: That persistent global summary is a really smart addition because it prevents the model from completely losing track of the overall scene context while it’s focusing on a specific detail.
Jane: So, in simple terms, Gaze Attention lets MLLMs perform more focused attention on task-relevant visual regions while simultaneously maintaining a compact, persistent global summary through those learnable context tokens.
The paper's improvements: Tom: Moving on to the specific improvements they propose in "Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal LLMs," the authors suggest a few key refinements to make this mechanism even better.
Jane: They point out that visual keys are partitioned into regions, where each region’s descriptor is obtained by mean-pooling the key vectors in Ru g, and then at decoding step j, the query measures its similarity to these descriptors to select the TopK regions.
Lu: That TopK selection mechanism is interesting because it allows the model to adapt their visual focus at each decoding step, rather than relying on a single relevance estimate from an earlier generation stage.
Meng: I’m interested in how they handle training stability; they mention that training visual routing from scratch can be unstable initially, so they adopted a progressive TopK schedule starting dense and gradually reducing K to encourage more stable routing behavior.
Lalam: That makes sense; it shows the authors are thinking about making the mechanism robust enough to actually learn effectively in practice, which is crucial for deployment.
Tom: And they also introduce the concept of learnable context tokens, which act as compact summary keys k uC, processed within the LLM’s native attention layers to provide access to that global summary.
Jane: So, it’s not just about how you select regions, but also how you integrate that persistent global context so the model doesn't get lost in the localized focus.
Conclusion: Tom: Alright, we’ve covered a lot about Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal LLMs; it seems like this paper successfully demonstrates how to move toward more focused and computationally efficient visual understanding in MLLMs.
Jane: To wrap up the implications, this work shows that selective attention mechanisms can yield performance matching dense-attention baselines while drastically reducing the visual KV cache by up to ninety percent.
Lu: The potential here is huge for creating models that can handle massive amounts of visual data efficiently, which opens up avenues for much more complex reasoning in multimodal tasks.
Meng: From an engineering standpoint, reducing the memory footprint so significantly means we can actually deploy these kinds of sophisticated visual models on devices that currently struggle with large vision-language models.
Lalam: I think this ability to maintain a concentrated view while retaining global awareness is really what will improve how we build AI systems for describing complex scenes; it mimics how we focus when describing something important.
Tom: So, Gaze Attention gives us a powerful tool that allows MLLMs to achieve more focused attention on task-relevant visual regions while still keeping a compact, persistent global summary through those learnable context tokens.
Jane: That’s the essence of it; it’s about smarter routing rather than brute-force attention across the board. We've explored how they achieve this and what that means for our future work in multimodal AI.
Junha Song, Byeongho Heo, Geonmo Gu, Jaegul Choo
KAIST · NAVER AI Lab
cs.CV
Submitted: 2026-05-13
Updated: 2026-10-02
Code: https://github.com/LLaVA-VL/LLaVA-NeXT2https:
Importance score: 90/100
The gist: MLLMs currently attend to all visual tokens during generation, leading to diluted focus and unnecessary computational overhead, whereas human visual perception is inherently selective.
Key concepts
- Gaze Regions
- Visual embeddings are spatially grouped into distinct areas called gaze regions. Each region is summarized by a lightweight descriptor derived from mean-pooling the key vectors within that group. This partitioning allows the model to treat specific visual areas as manageable, discrete units for routing attention.
- TopK Routing
- At each decoding step, the model calculates how relevant its current query is to these gaze region descriptors. It then selects only the 'TopK' most similar regions and restricts its attention solely to the visual key-value pairs within those selected regions. This dynamic selection enables focused attention on task-relevant parts of the image.
- Learnable Context Tokens
- To prevent losing global context when focusing locally, learnable context tokens are added to each image or frame. These tokens are processed by the LLM's standard attention layers and provide a persistent summary of the entire visual scene, ensuring that even when only a few regions are attended to, the model retains a compact overview.
- Progressive TopK Schedule
- Training is stabilized using a progressive TopK schedule. This schedule starts with a dense setting and gradually reduces the number of selected regions (K) as training progresses. This gradual reduction encourages more stable and reliable visual routing behavior during the learning phase.
Terminology
Summary
MLLMs currently attend to all visual tokens during generation, leading to diluted focus and unnecessary computational overhead, whereas human visual perception is inherently selective. This work introduces Gaze Attention, a novel mechanism that enables Multimodal Large Language Models (MLLMs) to selectively attend to task-relevant visual regions during generation by dynamically routing attention based on compact gaze regions and learnable context tokens.
How it works
-
Visual embeddings are spatially grouped into
gaze regions,
each represented by alightweight descriptor.
-
At each decoding step, the model dynamically selects the most relevant regions and restricts attention to them based on these descriptors, enabling the model to perform
more focused attention on task-relevant regions.
-
To maintain global context,
learnable context tokens
are appended to each image or frame and processed within the LLM’s native attention layers, providing apersistent global summary even when most visual regions are hidden.
Key Components and Mechanism
(Note: The paper enumerates these components as part of the method description.)
-
Visual keys are partitioned into regions, where each region is summarized by a descriptor obtained by
mean-pooling the key vectors in Ru g.
-
At decoding step j, the query qj measures its similarity to region descriptors and selects the
TopK
regions:s u g = q T j d u g, G j = TopK(s u g, K).
-
The model then attends only to the visual key–value pairs within these selected regions, allowing it to
adapt their visual focus at each decoding step.
-
To mitigate the loss of global context caused by localized attention, learnable context tokens are introduced, which are processed within the LLM’s native attention layers and provide access to a
compact summary keys k uC.
Training and Stability
(Note: The paper addresses training stability through specific scheduling.)
-
Training visual routing from scratch can be unstable because the relationship between descriptors and queries is weak initially.
-
A
progressive TopK schedule
is adopted, starting from a dense setting and gradually reducing K as training proceeds, which encouragesmore stable routing behavior.
Performance and Analysis
(Note: The paper demonstrates superior performance over baselines.)
-
Gaze Attention matches or surpasses dense-attention baselines across 13 image and 6 video understanding benchmarks.
-
The method reduces the visual KV cache by up to 90% while maintaining competitive performance, with computational cost scaling with the attended fraction rather than the full visual token count.
-
Visualization analyses confirm that Gaze Attention yields
concentrated, interpretable attention maps at each generation step,
contrasting withdiffuse patterns of dense attention.
-
In video understanding, Gaze Attention is hypothesized to be more effective than simply reducing the number of input tokens under a constrained visual KV budget.
Limitations and Future Directions
(Note: The paper outlines areas for future research.)
-
A limitation is the use of
Fixed-size regions,
which may not align well with semantic boundaries, suggesting thatContent-adaptive region partitioning could yield more reliable gaze routing.
-
The current descriptor design uses mean-pooling, treating all tokens within a region equally; future work could incorporate
entropy-based importance weighting
to improve descriptor quality. -
The model is trained solely with cross-entropy loss, lacking explicit supervision; a future direction is to investigate the effect of
explicit gaze guidance on visual routing,
potentially realized by combining sentence parsing with an open-vocabulary detector during training.
Computational Cost
(Note: The paper quantifies the efficiency gains.)
-
Gaze Attention reduces attention FLOPs by approximately 81% and KV-cache memory by 79% when utilizing the KV-cache offloading strategy of ReKV (Di et al., 2025; Ning et al., 2025).
-
With kernel-level algorithms like FlashAttention enabled, Gaze Attention incurs only a
2.6×
wall-clock time compared to dense baselines, which require 3.5× or 3.6× the time for existing methods under similar conditions.
The gist
Gaze Attention selectively attends to query-relevant visual regions by dynamically routing attention based on compact gaze regions and learnable context tokens, matching or surpassing dense-attention baselines while reducing the visual KV cache by up to 90%.
Improvements for AI systems
Here are the specific improvements for AI systems based on the Gaze Attention mechanism described in this paper:
-
Enhance selective visual attention in Multimodal Large Language Models (MLLMs) by replacing dense attention with a dynamic, query-adaptive routing system.
-
Implement a novel mechanism where visual embeddings are spatially grouped into compact
gaze regions,
each summarized by a lightweight descriptor. -
At every decoding step, dynamically select only the most relevant gaze regions based on these descriptors and restrict attention exclusively to them, thereby reducing redundant computation while maintaining high focus on task-relevant visual content.
-
Integrate learnable context tokens appended to each visual unit to provide a persistent, compact summary of global scene context within the LLM's native attention layers, ensuring holistic awareness alongside localized focus.
-
Enable models to adapt their visual focus dynamically during generation by re-selecting regions at each step (unlike KV-cache eviction methods that rely on a one-time importance estimate).
-
Improve efficiency by reducing the visual Key-Value (KV) cache entries attended to by up to 90% compared to dense attention baselines, leading to significant reductions in FLOPs and memory usage for both image and video understanding tasks.
-
Achieve superior performance on benchmarks (e.g., ImageQA, VideoQA) by matching or surpassing dense-attention baselines while exhibiting more concentrated, interpretable attention maps at each generation step.
-
Enable robust performance under aggressive KV budget constraints (e.g., attending to only 30-72 entries in single-image settings), mitigating the risk of insufficient visual information when using token compression or cache eviction methods on sparse inputs.
-
Develop more accurate and contextually grounded visual understanding, especially for video tasks, by allowing temporal routing across frames in addition to spatial region selection.
-
Create a system that mimics human cognitive behavior—selectively fixing gaze on relevant details while retaining a holistic scene view—enabling the AI to generate descriptions that are both highly focused on specific entities and aware of the broader context.
Sources
- GPT-4 Technical Report
- Longformer: The Long-Document Transformer
- Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs
- Mistral 7B
- Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction
- DeepGaze II: Reading fixations from deep features trained on object recognition
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
- LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
- Kimi-VL Technical Report
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Qwen2.5 Technical Report
- Qwen3 Technical Report
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models