MissClick: Execution-Aware Adversarial Attacks on Coordinate Generation in GUI Grounding Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "MissClick: Execution-Aware Adversarial Attacks on Coordinate Generation in GUI Grounding Models".
Jane: Recent GUI visual grounding models generate screen coordinates as sequences of digit tokens that are parsed into numerical values and mapped to executable clicks,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So we've covered a lot about "MissClick: Execution-Aware Adversarial Attacks on Coordinate Generation in GUI Grounding Models," and it seems clear that these models generate coordinates with a hidden structure that we can exploit by targeting how those numbers are formed. The authors propose MissClick-U for general disruption and MissClick-T for specific hijacking, showing how the place-value structure dictates the effectiveness of each attack objective.
Jane: It’s true, Tom; the paper highlights that simply looking at the coordinate string isn't enough because the actual click displacement depends entirely on those digit values and their decimal positions when parsed into numbers. The implications are that we have to move beyond treating these outputs as simple text and start considering their numerical construction when designing security for these systems.
Lu: I think the main contribution is formalizing this dependency between a high-order digit change and its resulting large coordinate shift, which provides the theoretical backbone for why these specific attack objectives are so effective in their respective settings. It’s a solid foundation for future work in adversarial analysis.
Meng: From a practical standpoint, it suggests that when we build systems reliant on visual grounding for actions, we need to ensure the underlying coordinate generation pipeline is as resistant to structural manipulation as possible, focusing on hardening those high-order positional inputs.
Lalam: I see the bigger picture here; this work pushes us to consider how AI interacts with the physical world through structured data and demands that our defenses evolve to understand that structure rather than just looking at surface-level outputs.
Tom: That’s a great way to put it, Lalam; so while we're excited about the technical details of MissClick, the real impact is forcing a necessary shift in how we think about securing these complex AI-driven interfaces moving forward.
Conclusion: Tom: So, we're wrapping up our deep dive into "MissClick: Execution-Aware Adversarial Attacks on Coordinate Generation in GUI Grounding Models," and I want to recap what this paper really boils down to for our listeners. It’s essentially about showing that how a computer translates a visual guess into an actual click is vulnerable because of the way coordinates are structured numerically.
Jane: Exactly, Tom; it explains that these grounding models produce text sequences that become numbers, and we can exploit the place-value system—like knowing a digit in the hundreds place matters much more than one in the ones place—to make clicks go where we want them.
Lu: I think what’s fascinating is how they formalized this dependency; they show a direct link between a high-order digit change and a massive shift in the executed coordinate. It opens up some really creative possibilities for how we can design systems that are sensitive to those structural weaknesses.
Meng: From an engineering standpoint, it tells me that if we're building applications that rely on these models for automation, we have to treat those coordinate outputs not just as arbitrary numbers but as a structured input where position truly matters.
Lalam: I see this work suggesting that the way AI processes and maps visual information is inherently tied to its numerical representation, which could fundamentally alter how we build trust in AI-driven interfaces across our entire culture.
Tom: It really puts a spotlight on the execution phase of an AI system, showing that we can attack the final output by manipulating the underlying math rather than just tricking the initial image recognition part.
Jane: And when you look at it from a simple perspective, it means that even if an AI is good at seeing what's on screen, we can still mess with its final action because of how those numbers are calculated internally.
Lu: The authors clearly set up two distinct attack paths, MissClick-U for general disruption and MissClick-T for specific hijacking, which shows a sophisticated understanding of the target model’s architecture.
Meng: I’m thinking about the practical impact on security: if this technique is understood, we can start building defenses that specifically check for these types of coordinate manipulation during runtime.
Lalam: It suggests a future where the security protocols for visual systems aren't just about input validation, but about understanding the entire mathematical journey from pixels to precise action.
Tom: Exactly! So, while we’ve covered how it works, this paper really hammers home that by attacking the place-value structure of coordinates, we can design defenses that are far more resilient against these types of manipulation.
Jane: And this whole concept is about making sure that the AI's understanding of its output isn't just a surface-level prediction but something we can actually control or at least predict within defined boundaries.
National University of Defense Technology
cs.AI
Submitted: 2026-08-04
Updated: 2026-09-28
Importance score: 74/100
The gist: Recent GUI visual grounding models generate screen coordinates as sequences of digit tokens that are parsed into numerical values and mapped to executable clicks, creating a security vulnerability
Key concepts
- Coordinate Structure
- GUI grounding models output coordinates as text sequences that are parsed using place values (like tens, hundreds). This structure means changing a single digit in a high-order position causes a much larger physical shift in the final executed click than changing a low-order digit.
- Soft Coordinate Displacement (MissClick-U)
- This is an untargeted attack that creates an adversarial example by calculating 'soft coordinates' based on expected digit values weighted by their place values. By minimizing the distance between this soft click and the true target, it maximizes the chance of pushing the final click far away from its intended location.
- Place-Weighted Target-Digit Loss (MissClick-T)
- For targeted attacks, this method optimizes discrete digit choices by using a loss function that weights errors based on place value. This forces the optimization to focus on high-order digits because an error at the hundreds place has a much greater impact on the final coordinate than an error at the ones place.
Terminology
Summary
Recent GUI visual grounding models generate screen coordinates as sequences of digit tokens that are parsed into numerical values and mapped to executable clicks, creating a security vulnerability that allows for large displacements in executed clicks based on the place-value structure of these coordinates. The gist: MissClick exploits the numerical and place-value structure of coordinate outputs by proposing two goal-specific objectives—soft-coordinate displacement for untargeted disruption and place-weighted target-digit optimization for targeted hijacking—to attack GUI grounding models.
Problem Formulation and Coordinate Structure
The paper frames the problem by considering a coordinate-generating GUI grounding model that autoregressively produces a textual answer sequence, which is then parsed into numerical coordinates. The core insight is that each position carries a decimal place value (×100, ×10, ×1), so high-order digit changes produce larger coordinate shifts.
This structure means a single-digit change at a high-order position therefore shifts the executed click by a much larger distance than the same change at a low-order position.
The paper formalizes this by defining how a parsed coordinate component is calculated from its constituent digits and their decimal place values, showing that the coordinate change depends on both the digit change and its decimal place value.
MissClick-U: Untargeted Attack Objective
The objective for an untargeted attack is to craft an adversarial example where the predicted click falls outside the ground-truth bounding box, defined as r(c) ∈/ Bgt.
Since directly differentiating the executed click from the input perturbation is non-differentiable due to argmax operations, MissClick-U constructs a continuous relaxation. It defines a soft coordinate
component by using expected digit values weighted by their place values:
˜dk,j = X 9 v=0 pk,j (v)
The soft coordinate is then formed as:
c˜k = X m k j=1 ak j ˜dk j.
The MissClick-U objective is then defined as the squared distance between this soft click and the ground-truth click u gt: L U = (r(˜c) − u gt) squared. This construction yields the differentiable path I → zk,j → pk,j → ˜dk,j → c˜ → r(˜c),
allowing for gradient-based optimization to maximize displacement.
MissClick-T: Targeted Attack Objective
For a targeted attack where the executed click must fall inside an attacker-specified region Btgt, MissClick-T focuses on optimizing the discrete digit selections toward the target. The paper notes that simply matching expected digit values does not guarantee success because matching an expected digit to the corresponding target value does not guarantee that the target digit is the argmax of pk j.
To address this, MissClick-T employs a place-weighted target-digit loss
(L T). This objective directly maximizes the probability assigned to the target digit rather than matching soft coordinate values. The loss is constructed by weighting per-digit cross-entropy by its place value:
LT = - X K k=1 X m k j=1 ak j log pk j (d tgt k,j) / (X K k=1 X m k j=1 ak j).
This weighting ensures that a one-unit digit error at the hundreds place changes the coordinate value by 100, whereas the same error at the ones place changes it by only 1,
thereby concentrating optimization on high-order positions.
Experimental Results and Objective Comparison
Experiments across two models (OS-Atlas and UGround) and three platforms demonstrate that both objectives are effective. In untargeted settings, MissClick-U achieves overall ASRs of 75.07% and 72.93%,
which is the highest among baselines, showing that soft-coordinate displacement yields the highest untargeted attack success rate.
Conversely, in targeted settings, Place-weighted digit CE achieves the highest ASR (44.86% and 62.67%
). The analysis confirms that Soft-coordinate distance is less effective than the digit-based objectives under the targeted setting,
suggesting that directly optimizing target-digit probabilities better matches the targeted attack goal.
Attack Optimization and Sensitivity Analysis
Both attacks utilize projected sign-gradient updates following PGD to optimize the perturbation δ, with MissClick-U using a "+ sign to maximize L U and MissClick-T using a
- sign to minimize L T. Furthermore, sensitivity analysis shows that
MissClick outperforms the corresponding baselines at every tested budget across different perturbation budgets. The iteration-budget analysis indicates that
MissClick requires fewer optimization iterations to reach the corresponding baseline performance and achieves stronger attack performance under limited iteration budgets.
Improvements for AI systems
Here are the specific improvements that can be made to current GUI grounding models, based on the MissClick framework:
-
Extend coordinate generation security analysis beyond treating coordinates as ordinary text by explicitly modeling their numerical structure (place-value system). This moves the security focus from simple token perturbations to exploiting how high-order digits influence coordinate displacement.
-
Implement two specialized adversarial attack objectives tailored to specific goals:
pinpoint and maximize soft-coordinate displacement
for untargeted disruption (MissClick-U) or minimize place-weighted target-digit loss
for targeted hijacking (MissClick-T).
-
Develop a differentiable mechanism to compute soft coordinates that explicitly incorporates the decimal place value of each digit, allowing gradients to accurately reflect the non-uniform impact of high vs. low-order digit changes on the final executed click position.
-
Enhance untargeted attack robustness by training against objectives that maximize the squared distance between a differentiable soft click and the ground-truth center point, rather than relying solely on non-differentiable token cross-entropy loss.
-
Improve targeted hijacking success rates by optimizing for target digit probabilities using a
place-weighted target-digit loss
(MissClick-T), which assigns higher penalties to errors in high-order digits (e.g., hundreds place) that cause larger coordinate shifts, thereby encouraging the model to select the correct high-impact digits.
These improvements will enable the improved AI system to:
-
Identify and exploit vulnerabilities in GUI grounding models by crafting perturbations that are guaranteed to move a predicted click outside a region (untargeted) or into a specific attacker-defined target area (targeted).
-
Achieve significantly higher success rates against existing baselines, as demonstrated by MissClick-U achieving up to 75% untargeted ASR on OS-Atlas and UGround.
-
Develop more resilient grounding models by training them using
MissClick
style adversarial objectives, forcing the model to be robust against attacks that exploit the numerical structure of its coordinate outputs rather than just superficial pixel-level noise.
Sources
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- Phi-Ground Tech Report: Advancing Perception in GUI Grounding
- Environmental Injection Attacks against GUI Agents in Realistic Dynamic Environments
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection