Benchmarking LLM Judges for Mobile Agent Evaluation
Ziqiang Wang, Li Gu, Zhixiang Chi, Zhi Liu, Seyed Mehdi Ayyoubzadeh, Yuanhao Yu, Yang Wang
Mila – Québec AI Institute · Concordia University · University of Toronto · Shanghai University · McMaster University
cs.AI, cs.CL, cs.CV
Submitted: 2026-08-11
Updated: 2026-08-17
Code: https://github.com/hiyouga/EasyR1
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: M OBILE J UDGE B ENCH is introduced as "to our knowledge the first benchmark for evaluating LLM-as-judge methods on mobile agent trajectories." The benchmark comprises "931 human-annotated
Terminology
Summary
M OBILE J UDGE B ENCH is introduced as to our knowledge the first benchmark for evaluating LLM-as-judge methods on mobile agent trajectories.
The benchmark comprises 931 human-annotated trajectories spanning 6 benchmarks, 4 agents, and 68 apps,
collected from 6 established mobile agent benchmarks, generated by 4 diverse agent models across 68 apps.
The paper evaluates 6 judge methods (five adapted from SPA-Bench, A3 with two modes, AndroidArena, and AgentRewardBench, plus a simple baseline we design) across multiple LLM backends.
The paper's key findings are threefold. First, "a simple baseline judge with sampled screenshots is competitive with, and often exceeds, purpose-built methods, indicating that more elaborate judge pipelines do not consistently improve judge quality; among competitive methods, the LLM backbone is the primary driver. Second,
benchmark quality metrics reliably predict real-world judge utility: they correlate with both agent ranking fidelity for evaluation and downstream performance when judges serve as reward signals for on-policy reinforcement learning. Third,
failure analysis across two LLM backends uncovers qualitatively opposite failure profiles, one conservative and the other permissive, linked to the backbone's precision-recall characteristics."
The paper's contributions are: (1) constructing a dataset of 931 human-annotated trajectories spanning 6 benchmarks, 4 agents, and 68 apps, together with a unified evaluation framework that standardizes judge assessment with classification and ranking metrics
; (2) systematically evaluating 6 judge methods across 5 LLM backends, revealing that no single method dominates and that judge accuracy depends substantially on both the method and the LLM backbone
; (3) designing a streamlined judge that is competitive with or exceeds purpose-built methods (up to 90.9% accuracy)
with ablations showing screenshot count is the dominant design variable, while the other inputs have marginal impact
; (4) validating the benchmark through meta-correlation analysis
showing judge quality metrics reliably predict agent ranking fidelity and success rate estimation accuracy
and on-policy RL experiments demonstrat[ing] that judge quality on our benchmark carries through to downstream agent performance
; and (5) analyzing "hard-core failure cases where nearly all judge methods err, revealing that different LLM backends produce qualitatively opposite failure profiles, one conservative (false-negative-heavy) and the other permissive (false-positive-heavy), with distinct root cause taxonomies."
The human annotation process involved 9 graduate-student annotators
labeling each trajectory with a binary success judgment (success/failure),
with Average pairwise agreement on the success label is 88.4%.
The final dataset is approximately balanced: 492 success (52.8%) and 439 failure (47.2%).
In the judge evaluation, no single method dominates: the simple baseline is the strongest method with the Gemini, GPT-5-mini, and Qwen backends (90.9%, 90.8%, 85.7%), while AgentRewardBench is strongest with GLM and Claude (86.1% and 87.2%).
The LLM backend substantially affects every method: the same method can vary by 5.6–10.5pp across backends.
A two-way variance decomposition over the 6×5 grid
shows method choice explains 49% of the accuracy variance and the backbone 21%,
but among AndroidArena, SPA-Bench, AgentRewardBench, and the baseline, backbone choice dominates
(11% vs. 49%).
The ablation study found "fewer higher-resolution screenshots outperform many low-resolution ones: 16 screenshots at 1/4 resolution achieves the best accuracy (91.2% for GPT, 86.5% for Qwen), while 192 at 1/64 degrades substantially (86.6%/82.3%). Also,
UI tree metadata has negligible impact (≤0.3pp for Qwen), and agent reasoning provides only a modest gain for GPT (∼1.7pp)."
For the meta-correlation validation, F1 is the strongest predictor of ranking fidelity (ρs = 0.90, 95% CI [0.66, 0.92]), while balanced accuracy best predicts rate estimation (ρs = −0.79 [−0.92, −0.70]).
Strikingly, precision has no predictive power for either metric (both of its CIs span zero); what matters is balanced classification, not precision alone.
The highest-accuracy judge (Baseline/Gemini, 90.9%) closely tracks human success rates (ρ = 0.97).
For the on-policy RL training experiments, judge accuracy predicts training outcomes: rule-based (94.6% acc) → 54.6% best easy-set success rate, GPT-5-mini (92.2%) → 45.4%, GPT-5.2 (88.8%) → 42.6%, Qwen (88.8%) → 39.9%.
The paper notes "GPT-5.2 and Qwen have identical accuracy but opposite precision–recall profiles, and the higher-precision GPT-5.2 (precision 93.7% vs. 80.2%) reaches a 2.7pp higher easy-set success rate, consistent with false positives directly rewarding incorrect behavior."
The failure analysis revealed "GPT-based judges produce 48 hard-core failures dominated by false negatives (30 FN, 18 FP): they are too conservative, failing to recognize successful trajectories. Qwen-based judges produce 78 hard-core failures dominated by false positives (71 FP, 7 FN): they are too permissive, accepting failed trajectories as successes. The root cause taxonomy shows
The GPT set is dominated by false negatives (last-frame anchoring, unfamiliar success state), while the Qwen set is dominated by false positives (constraint violation, partial completion). Surface UI match is a shared weakness across both backends. The classification was reliable with
two raters independently categorized a random sample of 30 hard-core failure cases... agreeing on 29 of 30 (Cohen’s κ = 0.957)."
The paper concludes that Judges should be evaluated, not assumed reliable,
Elaborate judge pipelines do not consistently outperform a simple baseline,
Different applications demand different judge profiles
(with F1 and balanced accuracy best predict reliability
for evaluation, while judges with strong false-positive control may be preferable as reward signals
), and Failure modes are structurally addressable
(e.g., Last-frame anchoring can be mitigated by providing more screenshots
).
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems:
-
Simplify judge pipelines: I will replace complex, purpose-built judge methods with a streamlined baseline that uses sampled screenshots, since it matches or exceeds elaborate methods (up to 90.9% accuracy) while reducing computational overhead.
-
Optimize screenshot sampling: I will use fewer, higher-resolution screenshots (e.g., 16 at 1/4 resolution) instead of many low-resolution ones, as this yields the best accuracy (91.2% for GPT, 86.5% for Qwen) and avoids degradation from excessive low-res frames.
-
Select LLM backends based on task profile: I will choose backends with high precision (e.g., GPT-5.2, precision 93.7%) for reward signals in RL training to avoid rewarding incorrect behavior, and backends with high balanced accuracy for evaluation tasks, since precision alone has no predictive power for ranking fidelity.
-
Adapt judge inputs to failure modes: I will provide more screenshots to mitigate
last-frame anchoring
(a false-negative cause in GPT-based judges) and add explicit constraint-checking prompts to reduce false positives fromconstraint violation
andpartial completion
(dominant in Qwen-based judges). -
Use benchmark metrics as selection criteria: I will use F1 score (ρs = 0.90) to predict agent ranking fidelity and balanced accuracy (ρs = −0.79) to predict success-rate estimation, allowing me to pre-select judges for specific evaluation or training purposes without costly real-world trials.
-
Implement dual-judge or ensemble systems: Since GPT-based judges are conservative (false-negative-heavy) and Qwen-based are permissive (false-positive-heavy), I will combine them or use a meta-judge that reconciles their outputs, reducing hard-core failures (48 vs. 78 cases) and improving overall reliability.
-
Incorporate UI tree metadata only when cheap: I will skip UI tree metadata for most cases (negligible impact ≤0.3pp) but include agent reasoning when using GPT backends (modest 1.7pp gain), to balance accuracy and cost.
-
Calibrate judge thresholds for RL: For on-policy RL, I will tune the judge’s decision threshold toward higher precision (as in GPT-5.2) to avoid false positives that directly reward incorrect actions, leading to higher downstream success rates (54.6% vs. 39.9% with lower-precision judges).
The improved AI system can: (a) evaluate mobile agent trajectories with near-human accuracy (90.9%) using a simple, fast baseline; (b) reliably rank agents and estimate success rates for benchmarking; (c) serve as an effective reward signal for RL training, improving agent performance by up to 14.7pp; and (d) adapt its failure profile (conservative vs. permissive) to the application, reducing systematic errors in both evaluation and training pipelines.
Sources
- Qwen2.5-VL Technical Report
- Constitutional AI: Harmlessness from AI Feedback
- A3: Android Agent Arena for Mobile GUI Agents with Essential-State Procedural Evaluation
- Agentic Reward Modeling: Verifying GUI Agent via Progressive Trajectory-Grounded Interaction
- ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
- The Llama 3 Herd of Models
- LLM Critics Help Catch LLM Bugs
- A Survey on LLM-as-a-Judge
- NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild
- The Dark Side of Rich Rewards: Understanding and Mitigating Noise in VLM Rewards
- Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
- Autonomous Evaluation and Refinement of Digital Agents
- Benchmarking Mobile Device Control Agents across Diverse Configurations
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- Preference Leakage: A Contamination Problem in LLM-as-a-judge
- LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
- An Illusion of Progress? Assessing the Current State of Web Agents
- HybridFlow: A Flexible and Efficient RLHF Framework
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection