Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets

arXiv:2606.25760 · cs.LG, cs.AI, cs.CL, cs.CV · Submitted 2026-06-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Uncertainty Quantification for Computer-Use Agents".

Jane: Computer-use agents transform vision-language model (VLM) predictions into executable GUI clicks, so reliable uncertainty estimates are essential for rejection, calibration, miss-severity ranking, and spatial safety regions.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we’re talking about "Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets," and the authors are looking at how UQ methods hold up when the agent or interface changes. It's a big question for anyone building systems that automate actions on screens.

Jane: Exactly, Tom; they are trying to find out if the rankings we see for uncertainty scores stay stable when you move from one VLM to another, or from one dataset to another, which is a crucial piece of information. They want to map out where these UQ methods generalize and where they fail.

Lu: The focus is on the fact that current evidence on post-hoc UQ for these agents is fragmented across isolated model and dataset pairs, leaving it unclear whether UQ method rankings stay stable when the agent, benchmark, or observable interface changes. They are addressing this fragmentation directly.

Meng: It sounds like they’re trying to create a standard way to test if a specific uncertainty metric is good enough for deployment in various real-world scenarios. That moves us past just picking the highest score on one paper and lets us see if that score actually means something when conditions shift.

Lalam: I think this is really important because if we can’t trust our uncertainty estimates across different contexts, then building truly robust computer-use agents becomes incredibly difficult because we can't predict how they will behave in novel situations.

The paper's summary: Tom: The core of the paper is introducing ARGUS, which is this cross-regime benchmark for post-hoc UQ in single-step executable GUI grounding. They evaluate a huge set of methods, including logit scores, sampling consistency measures like semantic entropy, and hidden-state probes.

Jane: They are looking at seven families of methods across four different VLM agents and datasets for the open-weight side, plus an eight-method matrix for closed-source APIs where things like logits aren't even available. This setup is designed to be as comprehensive as possible.

Lu: The main finding they highlight is something called selective transfer: UQ rankings are stable across datasets for a fixed model, but they degrade when you change the model class or the observable interface you are using. That’s a key distinction they make between different types of agent interactions.

Meng: That "selective transfer" idea suggests that we shouldn't treat all uncertainty scores as equal; their usefulness is tied to the specific setup of the agent and its environment. It implies that a method good for one type of model interaction might be useless in another, which is a very practical constraint.

Lalam: It’s fascinating because it shows that UQ quality isn't just about the math behind the score itself; it depends heavily on where you observe that score and how you intend to use it for rejection or calibration.

The paper's improvements: Tom: Regarding improvements, the authors point out that hidden-state and density methods form what they call the most stable open-weight family of UQ methods. They also noted that sampling-based scores and verbalised self-assessment win in specific regimes, which is a nice piece of targeted information.

Jane: They also found that ranking transfer is strongest when you keep the model fixed across different datasets, reaching a Spearman correlation of zero point nine six nine for open-weight pairs. This means if you stick with one model but test it on more data, the uncertainty scores are quite consistent.

Lu: Conversely, they found that cross-tier transfer to closed-source vendors is much weaker, with Spearman correlation averaging only +zero point zero eight over twelve vendor times dataset pairs on the shared eight-method intersection. That gap in performance across model classes and interfaces is a significant finding.

Meng: That weak transfer metric for closed-source APIs tells me that if we are relying on black-box APIs, we have to be much more careful about which specific UQ method we choose, because the results won't be as reliable when switching vendors.

Lalam: It reinforces the idea that UQ quality isn't just a property of the score; it depends entirely on where you observe and how you are using it, which is what they conclude.

Conclusion: Tom: So, to wrap up, this paper shows that UQ rankings aren't universal; they are stable within a fixed model across datasets but degrade when the model or interface changes. They also gave us a regime-aware selection recipe for choosing the right tool based on whether you need discrimination, calibration, or severity ranking.

Jane: It’s clear that deployment needs to be tailored to what you are trying to achieve; AUROC measures error detection differently than AUSE measures severity ranking, so you can't just use one score for everything.

Lu: The implication is that we need a smarter system that can dynamically classify the current operational regime—dataset, model class, interface type—and then select the UQ method proven most reliable for that specific context.

Meng: From an engineering standpoint, this means our deployment pipeline shouldn't use a fixed uncertainty score; it should be built with logic to switch its entire measurement strategy depending on whether it's dealing with internal model states or external API calls.

Lalam: This whole ARGUS benchmark is really valuable because it gives us a reproducible basis for making these regime-aware choices, which will help make those computer-use agents safer and more trustworthy in real applications.

University of Illinois Chicago

cs.LG, cs.AI, cs.CL, cs.CV

Submitted: 2026-06-24

Updated: 2026-09-29

Comments: Accepted at NeurIPS 2026. 32 pages, 3 figures, 26 tables

Journal ref: Advances in Neural Information Processing Systems (NeurIPS 2026)

Code: https://github.com/ENSTA-U2IS-AI/torch-uncertainty

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 87/100

The gist: Computer-use agents transform vision-language model (VLM) predictions into executable GUI clicks, so reliable uncertainty estimates are essential for rejection, calibration, miss-severity ranking,

Key concepts

Uncertainty Quantification (UQ)
UQ involves calculating how confident a computer agent is in its predictions. Since these agents perform actions like clicking on a screen based on visual input, knowing when the model is uncertain allows for better decision-making, such as rejecting unreliable clicks or prioritizing safer regions.
Selective Transfer
This describes how well different uncertainty scoring methods work when moved from one context to another. The paper found that UQ rankings are stable across different datasets if the underlying model remains the same, but they become less reliable when the model class or the visual interface changes.
Regime-Aware Selection
This is a strategy for choosing which uncertainty scoring method to use based on the specific situation. The paper suggests that you must select a method based on whether you have access to internal model signals (like hidden states) or only observable API responses, as this determines the best available UQ signal.
Deployment Objectives
The goal of using an uncertainty score dictates which type of UQ metric is most useful. For example, if the goal is to catch errors, you need a measure like AUROC; if the goal is to rank how severe a mistake is, you need AUSE; these objectives often lead to different preferred scoring methods.

Terminology

Summary

Computer-use agents transform vision-language model (VLM) predictions into executable GUI clicks, so reliable uncertainty estimates are essential for rejection, calibration, miss-severity ranking, and spatial safety regions. The main finding is selective transfer: UQ rankings are stable across datasets for a fixed model, but degrade across model classes and observable interfaces.

ARGUS Benchmark and Scope

The paper introduces ARGUS (Assessing Regime-wise Generalization of Uncertainty Scoring), a cross-regime benchmark for post-hoc UQ in single-step executable GUI grounding. It covers a "27-method, seven-family open-weight matrix over 4 GUI-grounding VLM agents and 4 datasets, plus an 8-method API-compatible closed-source matrix across 3 frontier vendors where logits, hidden states, and attention maps are unavailable." The evaluated methods span logit-based scores, sampling and consistency measures such as semantic entropy and self-consistency, hidden-state and density estimators such as Mahalanobis and SAPLMA, attention-based scores, P(True) and verbalised-confidence prompting, and split-conformal prediction.

Key Findings on UQ Generalization

The study investigates whether UQ methods generalize across regimes defined by the agent, dataset, and observable model interface. The main finding is selective transfer: UQ rankings are stable across datasets for a fixed model, but degrade across model classes and observable interfaces. Specifically:

Hidden-state and density methods form the most stable open-weight family,

CoCoA-1MCA, Focus, sampling-based scores, and verbalised self-assessment win in specific regimes.

The paper notes that Ranking transfer is strongest within a fixed model across datasets, reaching Spearman ρ = 0.969 and averaging ρ = 0.705 over 120 open-weight pairs. In contrast, cross-tier transfer to closed-source vendors is much weaker: Spearman ρ averages only +0.08 over 12 vendor×dataset pairs on the shared 8-method intersection.

Regime-Dependent Reliability and Transfer Mechanisms

The reliability of UQ families changes based on model class and interface observability. Model transitions show that attention, verbalised, and VLM-native families lose AUROC on every dataset, while density methods remain stable. Furthermore, API-only observability removes internal model signals (logits, hidden states, attention maps), promoting response-level signals such as Verbalised-1S and Verbalised-2S. The study concludes that UQ quality is not method-intrinsic; it depends on where the score is observed and how it is used.

Deployment Evaluation: Discrimination vs. Calibration

The paper evaluates UQ signals based on deployment objectives, finding that AUROC measures error detection, AUSE measures severity ranking, and calibration measures risk interpretability. For instance, AUROC and AUSE winners disagree on 14 of 16 open-weight cells, indicating that binary error detection and graded miss-severity ranking often prefer different UQ signals. Additionally, for spatial deployment, "conformal click regions show that score-level discrimination is not enough for deployment: locally weighted disks can shrink radii by 40–60% when the plug-in UQ is calibrated, but coverage can degrade under calibration-test or interface mismatch."

Practical Selection Strategy

The paper provides a Regime-aware UQ selection recipe to guide deployment. The protocol suggests:

  1. If hidden states are available, include SAPLMA / SEP and Mahal-RMD. If API-only, start with the harmonised 8-method panel: CCP, SelfCons, SE, LexSim, Verb-1S, Verb-2S, HEDGE, and IMGHEDGE.

  2. Density/probe is the most stable openweight family; API-only regimes promote response-level and verbalised scores.

  3. For severity ranking (graded severity), validate AUSE; for calibrated risk, check ECE / Brier after isotonic calibration.

  4. Reuse a prior panel across datasets only when the model is fixed; if model family or interface changes, rerank on the target calibration split.

  5. For spatial coverage, start with Disk-Fixed or Disk-CQR; use Disk-Normalized only after coverage checks pass. The final recommendation is that UQ selection should be regime-specific.

Release and Reproducibility

The authors release ARGUS as a Python package named argus-uq, including 27 method implementations, per-item records, splits, UQ scores, closed-source API responses, and an analysis pipeline. This provides a reproducible basis for regime-aware UQ selection in GUI agents.

Improvements for AI systems

Based on the scientific paper Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets, here are specific, actionable improvements you can implement in AI systems, categorized by their function:


)1. Robust Decision Making and Rejection Strategy (Risk Management):

The system should move beyond simple confidence scores to implement a multi-objective risk assessment pipeline informed by the ARGUS benchmark findings:

  • Implement a dynamic panel selection protocol based on the current context (model, dataset, interface). If using an open-weight model and switching datasets/interfaces, the system should automatically switch its UQ panel from density/probe methods to verbalised self-assessment methods to maintain reliability.

  • For high-stakes actions (e.g., destructive file operations), prioritize metrics that measure calibrated risk (ECEiso, Brieriso) over raw error detection (AUROCincorrect). The system should only proceed if a specific UQ method admits a non-trivial calibration threshold at the deployment confidence level (as suggested by CRC analysis).

  • Implement selective execution logic: if the uncertainty score exceeds a calibrated threshold, the system should not execute; instead, it should trigger an alternative action like deferral, logging the event for human review.

)2. Enhanced Error Discrimination and Severity Ranking:

The system needs to distinguish between wrong clicks and badly wrong clicks:

  • Integrate graded severity ranking (using AUSE on log(1 + dnorm)) into the decision loop. This allows the system to rank potential errors by their magnitude, enabling more nuanced responses than binary correctness alone.

  • When an error is detected, the system should use a UQ method that excels at identifying severe errors (like certain density/probe methods) to prioritize immediate rejection or high-level intervention over simple deferral.

)3. Adaptive Spatial Safety Regions:

Instead of using fixed confidence thresholds for click regions, the system should employ conformal prediction techniques:

  • Implement Disk-Normalized or Disk-CQR uncertainty measures to define spatial safety disks around predicted coordinates. These methods are shown to shrink radii by 40–60% when calibrated, leading to more precise and less over-covering safety regions compared to fixed-radius disks.

  • The system should continuously monitor the coverage of these adaptive disks against the target bounding box, dynamically adjusting its spatial uncertainty based on real-time calibration data from the deployment environment.

)4. Model/Interface Transition Robustness:

The system must be resilient when moving between different model architectures or API access levels:

  • When transitioning from a vanilla VLM (e.g., Q72) to a grounding specialist (e.g., UI-TARS), the system should automatically re-evaluate its preferred UQ family, shifting reliance away from attention and verbalized families toward density/probe methods, which are shown to be more stable across model class transitions.

  • When moving from open-weight internals to closed-source API usage (losing logits/hidden states), the system must switch its reliance entirely to the harmonized 8-method panel (CCP, SelfCons, SE, LexSim, Verb-1S/2S, HEDGE, IMGHEDGE) and explicitly avoid extrapolating recommendations from open-weight proxies.

)5. Automated UQ Method Selection (Regime-Awareness):

The system should not rely on a single UQ score but use the ARGUS framework to select the optimal method for the current task:

  • Develop a meta-layer that classifies the current operational regime (Dataset, Model Class, Interface Type). Based on this classification, it selects the UQ method proven most reliable in that specific cell (e.g., Density/Probe methods for OSWORLD-G grounding vs. Verbalised scores for Gemini API calls).

)6. Post-Deployment Verification and Iteration:

The system needs a mechanism to validate its UQ choices against the actual deployment environment:

  • Implement a feedback loop where the system compares its predicted uncertainty (e.g., calculated disk radius) with observed outcomes (coverage/misses). If coverage degrades due to an interface mismatch or calibration shift, the system must flag this regime change and trigger a reranking of its UQ panel for the next inference cycle.

Abstract

Computer-use agents turn vision-language model (VLM) predictions into executable GUI clicks, so reliable uncertainty estimates are essential for rejection, calibration, miss-severity ranking, and spatial safety regions. Yet evidence on post-hoc uncertainty quantification (UQ) for these agents is fragmented across isolated model and dataset pairs, leaving it unclear whether UQ rankings stay stable when the agent, benchmark, or observable interface changes. We present Argus, a cross-regime benchmark for post-hoc UQ in single-step executable GUI grounding: a 27-method open-weight matrix over 4 VLM agents and 4 datasets, plus an 8-method closed-source matrix across 3 frontier vendors where logits, hidden states, and attention maps are unavailable. Evaluated methods span logit-based scores, sampling and consistency measures, hidden-state and density estimators (Mahalanobis, SAPLMA), attention-based scores, P(True) and verbalised-confidence prompting, and split-conformal prediction. The main finding is selective transfer: UQ rankings are stable across datasets for a fixed model, but degrade across model classes and observable interfaces. Hidden-state and density methods are the most stable open-weight family, while CoCoA-1MCA, Focus, sampling-based scores, and verbalised self-assessment win in specific regimes. Within-model ranking transfer is strong (Spearman rho up to 0.969), but cross-tier transfer to closed-source vendors averages only +0.08, so closed-source UQ should be reranked on the target rather than extrapolated. Conformal click regions show score-level discrimination is not enough for deployment: locally weighted disks shrink radii by 40-60% when the plug-in UQ is calibrated, but coverage degrades under calibration-test or interface mismatch. We release per-item records, calibration/test splits, UQ scores, and analysis scripts for regime-aware UQ selection in GUI agents.

Sources

Related papers