GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness".
Jane: The gist The GUI-KV method introduces a plug-and-play KV cache compression technique for GUI agents that exploits spatio-temporal redundancy to achieve near–full-cache accuracy with modest budgets and significant reductions in…
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, moving on to the title and authors for "GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness," we have Kung-Hsiang Huang, Haoyi Qiu, Yutong Dai, Caiming Xiong, and Chien-Sheng Wu as the researchers behind this work.
Jane: They are from Salesforce AI Research at UC Los Angeles and they are focused on making these agents more efficient by focusing specifically on GUI tasks.
Lu: The title itself really sets the stage for what’s happening here—it’s about efficiency through KV cache management, but they added a crucial layer: spatio-temporal awareness.
Meng: Spatio-temporal awareness sounds like they are looking at both where things are spatially on the screen and how things change over time.
Tom: Right, and that’s the core idea. It means their method is designed to exploit the specific redundancies found in graphical user interfaces, which is different from general vision tasks.
Jane: This paper essentially says we need to stop treating GUI agents like they are just standard vision-language models and start accounting for their unique visual structure.
Lu: They motivate this because they first analyzed attention patterns in GUI agent workloads and found that the sparsity is actually uniformly high across all transformer layers.
Meng: That's a pretty strong finding, suggesting that the complexity of layer-varying strategies isn't necessary if we use a simple uniform budget allocation.
Tom: They show empirically that this simple uniform strategy outperforms more complex layer-varying schemes because of how flat and extreme the sparsity is in GUIs.
Jane: So, they are pushing back against existing methods that might be over-amplifying small differences or misallocating the cache based on those patterns.
Lu: They argue that existing KV cache approaches developed for natural images and documents aren't optimal because they don't account for the spatial and temporal redundancies specific to GUIs.
Meng: It’s a good reminder that you can’t just take a general technique and apply it universally without checking if the data structure matches.
Tom: And that’s what GUI-KV is designed to fix, making it a plug-and-play compression method for GUI agents that doesn't require any retraining.
Jane: So they are giving us a ready-to-use solution for this memory problem in agent workflows without needing to retrain the underlying models.
The paper's summary: Tom: Now let’s get into the actual summary of "GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness." They introduce GUI-KV as a plug-and-play compression method that exploits spatio-temporal redundancy.
Jane: In simple terms, they are taking two main techniques: spatial saliency guidance and temporal redundancy scoring to selectively keep tokens.
Lu: The spatial guidance uses the L2 norm of hidden states to preserve semantically important visual tokens, defined by a linear combination of attention and the norm of hidden states for visual and text tokens.
Meng: So it’s prioritizing the visual meaning over just what the model is looking at at that exact moment. That sounds like a smart way to keep the "what" and "how" connected.
Tom: And then they have temporal redundancy scoring, which uses QR decomposition over current screenshots to project previous frames' keys onto the current frame’s key subspace to prune redundant history.
Jane: This is basically making sure that if a visual element in a previous screenshot is just repeating itself in the current one, we don't need to keep it cached.
Lu: The final scoring function then combines both signals, and they select the top-scoring tokens based on a uniform budget applied across all heads.
Meng: And the selection process uses a redundancy threshold based on the percentile of residual norms to define what counts as temporally unique information.
Tom: So, they are ensuring that whatever makes it into the cache is both intrinsically important and temporally unique, which is their key mechanism for token selection.
Jane: It’s about filtering out the noise—the stuff that isn't worth storing in memory for these specific image interactions.
The paper's improvements: Tom: Let's talk about the specific improvements they propose in this paper. Improvement number one is adopting a uniform budget strategy because of that attention sparsity finding.
Jane: That means they are explicitly telling developers to stop using complex layer-varying schemes and just use an equal budget for every layer for GUI agents.
Lu: Second improvement is introducing residual-stream saliency guidance, which augments attention scores with the L2 norm of hidden states to better preserve semantically important visual tokens.
Meng: So they are making sure the importance of a visual token isn't just dependent on its attention score, but also on its overall magnitude in the hidden state.
Tom: And third improvement is implementing temporal redundancy scoring by performing QR decomposition over screenshots to determine redundancy for past visual tokens.
Jane: That’s a very concrete engineering step—they are using a QR rank of thirty-two for their projector because it spans a compact representation of the latest frame.
Lu: Finally, they enhance token selection by defining a redundancy threshold based on the percentile of residual norms to create that final score ψˆh.
Meng: So this method gives us a clear pipeline: spatial guidance and temporal redundancy work together to select tokens that are both important and unique.
Tom: And the result is that when they deploy it, they achieve a reduction in decoding FLOPs by thirty-eight point nine percent while increasing step accuracy by four point one percent over the full-cache baseline on AgentNetBench with five screenshots.
Jane: That specific number is what really grounds this paper—it shows measurable efficiency gains in a real test scenario, not just theoretical gains on paper.
Conclusion: Tom: So to wrap up the "GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness," they successfully show how to exploit spatio-temporal structure through spatial saliency guidance and temporal redundancy scoring.
Jane: It’s a successful method because it proves that GUI agents can operate effectively with significantly reduced memory while maintaining high performance across various benchmarks.
Lu: The main implication is that prior studies on budget allocation strategies for natural images are not transferable to GUI agent tasks.
Meng: So the big picture is that we need specialized techniques tailored to the unique visual nature of GUIs if we want these agents to be truly efficient.
Lalam: From a model perspective, this means the advances can improve culture because they allow us to deploy these models in more contexts where memory constraints are previously prohibitive.
Tom: That’s a powerful point from Lalam—moving from theoretical potential to practical deployment benefits for the end user experience.
Jane: Exactly. The paper concludes that GUI-KV recovers near–full-cache accuracy at modest budgets and consistently outperforms strong baselines over a wide range of compression ratios.
Lu: So the final result is that combining spatial saliency guidance and temporal redundancy scoring ensures that we retain tokens that are both intrinsically important and temporally unique.
Meng: It’s a focused approach, ensuring you keep tokens that are both visually relevant and temporally distinct from what came before.
Salesforce AI Research
cs.CL
Submitted: 2025-10-01
Updated: 2026-10-07
Importance score: 90/100
The gist: The gist The GUI-KV method introduces a plug-and-play KV cache compression technique for GUI agents that exploits spatio-temporal redundancy to achieve near–full-cache accuracy with modest budgets
Key concepts
- Attention Sparsity
- This refers to how much of the attention mechanism is actually used across different parts of a transformer model. The paper found that in GUI workloads, this sparsity is uniformly very high across all layers, suggesting a simple budget allocation strategy works best.
- Spatial Saliency Guidance
- This technique uses the L2 norm of hidden states to score visual tokens. It helps the system prioritize and preserve visually important elements within the screenshots by combining attention scores with this norm, ensuring semantically critical visual information is kept.
- Temporal Redundancy Scoring
- This mechanism identifies and prunes redundant history by projecting previous frames' keys onto the current frame's key subspace. By performing a QR decomposition on current screenshots, it determines which visual information from prior steps is repeated and can be safely discarded.
- Uniform Budget Allocation Strategy
- The authors found that assigning an equal KV cache budget to every transformer layer performs better for GUI agents than complex, layer-specific schemes. This simplicity is motivated by the uniform high attention sparsity observed in GUI workloads.
Terminology
Summary
The gist The GUI-KV method introduces a plug-and-play KV cache compression technique for GUI agents that exploits spatio-temporal redundancy to achieve near–full-cache accuracy with modest budgets and significant reductions in decoding FLOPs.
Analysis of Attention Patterns and Budget Allocation
The paper first analyzes attention patterns in GUI agent workloads and finds that attention sparsity is uniformly high across all transformer layers
. This insight motivates a simple uniform budget allocation strategy, which we show empirically outperforms more complex layer-varying schemes
. The analysis indicates that "GUI screenshots exhibit uniformly extreme and flat attention sparsity across layers (> 0.99), whereas natural images can have sparsity reaches as low as 0.88. This finding suggests that existing budget allocation strategies may be suboptimal for GUI agents because methods like PyramidKV and VL-Cache
over-amplify tiny differences and misallocate cache. Therefore, the authors posit that
it is best to assign uniform KV cache budget to each layer for GUI agents".
GUI-KV Compression Mechanism
GUI-KV combines two novel techniques for compression. First, it introduces spatial saliency guidance,
which uses the L2 norm of hidden states to better preserve semantically important visual tokens
. This is formalized by a scoring function where the score is a linear combination of attention and norm: ψh i = (Ah i + αSi if i ∈ It (visual tokens) A h i if i ∈ Tt (text tokens). Second, it introduces temporal redundancy scoring,
which projects previous frames’ keys onto the current frame’s key subspace to preferentially prune redundant history
. This temporal redundancy scoring involves a QR decomposition over the current screenshots to determine redundant visual information from prior steps.
Synergistic Scoring and Token Selection
The two mechanisms work synergistically to determine the final retention set. The final scoring function combines both signals: ψˆh i = (Ah i + αSi if i ∈ It (current frame visual tokens) (A h i + αSi) · 1[ρ h i ≥ τ h] if i ∈ I<t (previous frames visual tokens) A h i if i ∈ T (text tokens). The process then selects the top-scoring tokens based on a budget γ applied uniformly across all heads. This selection is done by choosing the indices with the top scores based on ψˆh i, where k = ⌈n · γ⌉ is the number of tokens to keep based on budget.
Experimental Results and Efficiency Gains
GUI-KV demonstrates superior performance across six benchmarks, recovering near–full-cache accuracy with modest budgets (typically 10–20%)
. For the UI-TARS-1.5-7B model on the AgentNetBench benchmark with five screenshots and a 40% budget, GUI-KV reduces decoding FLOPs by 38.9% while increasing step accuracy by 4.1% over the full-cache baseline
. The pre-filling overhead is negligible (increase < 0.29%), while decoding compute drops substantially, with GUI-KV 20 123.6 (-42.0%)
MFLOPs/ decoded token on AgentNetBench.
Ablation and Component Effectiveness
Ablation studies confirm the complementary benefits of the components. The spatial-only guidance is more effective with fewer screenshots (≤ 7), while temporal-only guidance excels as more screenshots are added, capitalizing on cross-frame redundancy. Combining both guidance results in the best performance across settings. Furthermore, the L2 norm of the residual stream is shown to be most effective
among different spatial saliency guidance approaches. The QR rank r = 32 was set for the projector in §3.2 because it spans a compact representation of the latest frame
.
Conclusion
GUI-KV successfully exploits the spatio-temporal structure of GUI interactions by combining spatial saliency guidance with temporal redundancy scoring to selectively retain informative tokens. This approach enables GUI agents to operate with significantly reduced memory while maintaining high performance across various benchmarks. The method proves that exploiting GUI-specific redundancies enables efficient and reliable agent performance. This work demonstrates that prior studies on budget allocation strategies developed for natural images are not transferable to GUI agent tasks. The paper concludes by showing that GUI-KV recovers near–full-cache accuracy at modest budgets and consistently outperforms strong baselines over a wide range of compression ratios. The efficiency gains are substantial, with decoding compute dropping substantially with negligible prefill overhead. The final result is that the combination of spatial saliency guidance and temporal redundancy scoring ensures that we retain tokens that are both intrinsically important and temporally unique.
--- Page 1 ---
GUI-KV: EFFICIENT GUI AGENTS VIA KV CACHE WITH SPATIO-TEMPORAL AWARENESS
WITH SPATIO-TEMPORAL AWARENESS
Kung-Hsiang Huang1 Haoyi Qiu2∗ Yutong Dai1 Caiming Xiong1 Chien-Sheng Wu1
Salesforce AI Research 2University of California, Los Angeles
ABSTRACT
Graphical user interface (GUI) agents built on vision–language models have emerged as a promising approach to automate human-computer workflows. However, they also face the inefficiency challenge as they process long sequences of high-resolution screenshots and solving long-horizon tasks, making inference slow, costly and memory-bound. While key-value (KV) caching can mitigate this, storing the full cache is prohibitive for image-heavy contexts. Existing cache-compression methods are sub-optimal as they do not account for the spatial and temporal redundancy of GUIs. In this work, we first analyze attention patterns in GUI agent workloads and find that, unlike in natural images, attention sparsity is uniformly high across all transformer layers. This insight motivates a simple uniform budget allocation strategy, which we show empirically outperforms more complex layer-varying schemes. Building on this, we introduce GUI-KV, a plug-and-play KV cache compression method for GUI agents that requires no retraining.
Improvements for AI systems
-
Improve KV cache allocation by adopting a uniform budget strategy; specifically,
we posit that it is best to assign uniform KV cache budget to each layer for GUI agents
because existing approaches fail to capture theuniformly extreme and flat attention sparsity across all layers
in GUIs, as shown in Figure 1. -
Introduce residual-stream saliency guidance by augmenting attention scores with the
L2 norm of hidden states
to better preserve semantically important visual tokens, as defined in Equation (4), which is then combined additively with attention scores to form the final selection score ψh. -
Implement temporal redundancy scoring by performing a
QR decomposition over the current screenshots
and projecting previous frame keys onto the current frame’s key subspace, as described in Equation (6) and (7), to quantify non-redundancy for past visual tokens. -
Enhance token selection by defining a redundancy threshold τh based on the percentile of residual norms, "τh ← Percentile1−γ
5:
The final scoring function combines both spatial saliency and temporal redundancy, selecting the top-k tokens based on scores ψˆh, which ensures they are both intrinsically important and temporally unique.
- Deploy this GUI-KV method to GUI agents to achieve a reduction in decoding FLOPs by 38.9% while increasing step accuracy by 4.1% over the full-cache baseline when using five screenshots on the AgentNetBench benchmark, as noted in Table 2 and Section 4.3.
Sources
- Less is More: Empowering GUI Agent with Context-Aware Simplification
- MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
- GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
- ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use
- Foundations of GenIR
- LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
- D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models
- OpenCUA: Open Foundations for Computer-Use Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering