Benchmarking LLM Judges for Mobile Agent Evaluation

arXiv:2608.11434 · cs.AI, cs.CL, cs.CV · Submitted 2026-08-11 · Read on arXiv

Ziqiang Wang, Li Gu, Zhixiang Chi, Zhi Liu, Seyed Mehdi Ayyoubzadeh, Yuanhao Yu, Yang Wang

Mila – Québec AI Institute · Concordia University · University of Toronto · Shanghai University · McMaster University

cs.AI, cs.CL, cs.CV

Submitted: 2026-08-11

Updated: 2026-08-17

Code: https://github.com/hiyouga/EasyR1

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: M OBILE J UDGE B ENCH is introduced as "to our knowledge the first benchmark for evaluating LLM-as-judge methods on mobile agent trajectories." The benchmark comprises "931 human-annotated

Terminology

Summary

M OBILE J UDGE B ENCH is introduced as to our knowledge the first benchmark for evaluating LLM-as-judge methods on mobile agent trajectories. The benchmark comprises 931 human-annotated trajectories spanning 6 benchmarks, 4 agents, and 68 apps, collected from 6 established mobile agent benchmarks, generated by 4 diverse agent models across 68 apps. The paper evaluates 6 judge methods (five adapted from SPA-Bench, A3 with two modes, AndroidArena, and AgentRewardBench, plus a simple baseline we design) across multiple LLM backends.

The paper's key findings are threefold. First, "a simple baseline judge with sampled screenshots is competitive with, and often exceeds, purpose-built methods, indicating that more elaborate judge pipelines do not consistently improve judge quality; among competitive methods, the LLM backbone is the primary driver. Second, benchmark quality metrics reliably predict real-world judge utility: they correlate with both agent ranking fidelity for evaluation and downstream performance when judges serve as reward signals for on-policy reinforcement learning. Third, failure analysis across two LLM backends uncovers qualitatively opposite failure profiles, one conservative and the other permissive, linked to the backbone's precision-recall characteristics."

The paper's contributions are: (1) constructing a dataset of 931 human-annotated trajectories spanning 6 benchmarks, 4 agents, and 68 apps, together with a unified evaluation framework that standardizes judge assessment with classification and ranking metrics; (2) systematically evaluating 6 judge methods across 5 LLM backends, revealing that no single method dominates and that judge accuracy depends substantially on both the method and the LLM backbone; (3) designing a streamlined judge that is competitive with or exceeds purpose-built methods (up to 90.9% accuracy) with ablations showing screenshot count is the dominant design variable, while the other inputs have marginal impact; (4) validating the benchmark through meta-correlation analysis showing judge quality metrics reliably predict agent ranking fidelity and success rate estimation accuracy and on-policy RL experiments demonstrat[ing] that judge quality on our benchmark carries through to downstream agent performance; and (5) analyzing "hard-core failure cases where nearly all judge methods err, revealing that different LLM backends produce qualitatively opposite failure profiles, one conservative (false-negative-heavy) and the other permissive (false-positive-heavy), with distinct root cause taxonomies."

The human annotation process involved 9 graduate-student annotators labeling each trajectory with a binary success judgment (success/failure), with Average pairwise agreement on the success label is 88.4%. The final dataset is approximately balanced: 492 success (52.8%) and 439 failure (47.2%).

In the judge evaluation, no single method dominates: the simple baseline is the strongest method with the Gemini, GPT-5-mini, and Qwen backends (90.9%, 90.8%, 85.7%), while AgentRewardBench is strongest with GLM and Claude (86.1% and 87.2%). The LLM backend substantially affects every method: the same method can vary by 5.6–10.5pp across backends. A two-way variance decomposition over the 6×5 grid shows method choice explains 49% of the accuracy variance and the backbone 21%, but among AndroidArena, SPA-Bench, AgentRewardBench, and the baseline, backbone choice dominates (11% vs. 49%).

The ablation study found "fewer higher-resolution screenshots outperform many low-resolution ones: 16 screenshots at 1/4 resolution achieves the best accuracy (91.2% for GPT, 86.5% for Qwen), while 192 at 1/64 degrades substantially (86.6%/82.3%). Also, UI tree metadata has negligible impact (≤0.3pp for Qwen), and agent reasoning provides only a modest gain for GPT (∼1.7pp)."

For the meta-correlation validation, F1 is the strongest predictor of ranking fidelity (ρs = 0.90, 95% CI [0.66, 0.92]), while balanced accuracy best predicts rate estimation (ρs = −0.79 [−0.92, −0.70]). Strikingly, precision has no predictive power for either metric (both of its CIs span zero); what matters is balanced classification, not precision alone. The highest-accuracy judge (Baseline/Gemini, 90.9%) closely tracks human success rates (ρ = 0.97).

For the on-policy RL training experiments, judge accuracy predicts training outcomes: rule-based (94.6% acc) → 54.6% best easy-set success rate, GPT-5-mini (92.2%) → 45.4%, GPT-5.2 (88.8%) → 42.6%, Qwen (88.8%) → 39.9%. The paper notes "GPT-5.2 and Qwen have identical accuracy but opposite precision–recall profiles, and the higher-precision GPT-5.2 (precision 93.7% vs. 80.2%) reaches a 2.7pp higher easy-set success rate, consistent with false positives directly rewarding incorrect behavior."

The failure analysis revealed "GPT-based judges produce 48 hard-core failures dominated by false negatives (30 FN, 18 FP): they are too conservative, failing to recognize successful trajectories. Qwen-based judges produce 78 hard-core failures dominated by false positives (71 FP, 7 FN): they are too permissive, accepting failed trajectories as successes. The root cause taxonomy shows The GPT set is dominated by false negatives (last-frame anchoring, unfamiliar success state), while the Qwen set is dominated by false positives (constraint violation, partial completion). Surface UI match is a shared weakness across both backends. The classification was reliable with two raters independently categorized a random sample of 30 hard-core failure cases... agreeing on 29 of 30 (Cohen’s κ = 0.957)."

The paper concludes that Judges should be evaluated, not assumed reliable, Elaborate judge pipelines do not consistently outperform a simple baseline, Different applications demand different judge profiles (with F1 and balanced accuracy best predict reliability for evaluation, while judges with strong false-positive control may be preferable as reward signals), and Failure modes are structurally addressable (e.g., Last-frame anchoring can be mitigated by providing more screenshots).

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems:

  1. Simplify judge pipelines: I will replace complex, purpose-built judge methods with a streamlined baseline that uses sampled screenshots, since it matches or exceeds elaborate methods (up to 90.9% accuracy) while reducing computational overhead.

  2. Optimize screenshot sampling: I will use fewer, higher-resolution screenshots (e.g., 16 at 1/4 resolution) instead of many low-resolution ones, as this yields the best accuracy (91.2% for GPT, 86.5% for Qwen) and avoids degradation from excessive low-res frames.

  3. Select LLM backends based on task profile: I will choose backends with high precision (e.g., GPT-5.2, precision 93.7%) for reward signals in RL training to avoid rewarding incorrect behavior, and backends with high balanced accuracy for evaluation tasks, since precision alone has no predictive power for ranking fidelity.

  4. Adapt judge inputs to failure modes: I will provide more screenshots to mitigate last-frame anchoring (a false-negative cause in GPT-based judges) and add explicit constraint-checking prompts to reduce false positives from constraint violation and partial completion (dominant in Qwen-based judges).

  5. Use benchmark metrics as selection criteria: I will use F1 score (ρs = 0.90) to predict agent ranking fidelity and balanced accuracy (ρs = −0.79) to predict success-rate estimation, allowing me to pre-select judges for specific evaluation or training purposes without costly real-world trials.

  6. Implement dual-judge or ensemble systems: Since GPT-based judges are conservative (false-negative-heavy) and Qwen-based are permissive (false-positive-heavy), I will combine them or use a meta-judge that reconciles their outputs, reducing hard-core failures (48 vs. 78 cases) and improving overall reliability.

  7. Incorporate UI tree metadata only when cheap: I will skip UI tree metadata for most cases (negligible impact ≤0.3pp) but include agent reasoning when using GPT backends (modest 1.7pp gain), to balance accuracy and cost.

  8. Calibrate judge thresholds for RL: For on-policy RL, I will tune the judge’s decision threshold toward higher precision (as in GPT-5.2) to avoid false positives that directly reward incorrect actions, leading to higher downstream success rates (54.6% vs. 39.9% with lower-precision judges).

The improved AI system can: (a) evaluate mobile agent trajectories with near-human accuracy (90.9%) using a simple, fast baseline; (b) reliably rank agents and estimate success rates for benchmarking; (c) serve as an effective reward signal for RL training, improving agent performance by up to 14.7pp; and (d) adapt its failure profile (conservative vs. permissive) to the application, reducing systematic errors in both evaluation and training pipelines.

Sources

Related papers