Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation

arXiv:2608.11191 · cs.CV, cs.AI, cs.CL · Submitted 2026-08-11 · Read on arXiv

Shiyu Xuan, Zechao Li

Nanjing University of Science and Technology

cs.CV, cs.AI, cs.CL

Submitted: 2026-08-11

Updated: 2026-08-12

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper introduces a Test-Time Self-Evolving framework for GUI visual grounding that enables models to improve after deployment without human-annotated ground truth.

Terminology

Summary

This paper introduces a Test-Time Self-Evolving framework for GUI visual grounding that enables models to improve after deployment without human-annotated ground truth. The framework constructs a closed-loop of Exploration, Evaluation, Reflection, and Internalization. Specifically, the agent first explores unseen interfaces by predicting grounding coordinates for given instructions. To evaluate these explorations, the authors introduce an MLLM-based Reflector to assess the generated results and provide corresponding reasoning reflections. To internalize reflection knowledge into the model weights, they propose Reflection-Guided On-Policy Self-Distillation (R-OPSD), which translates high-level reasoning into dense token-level supervision via a conditioned self-teacher. Furthermore, they design a Contrastive Calibration method to prevent incorrect auto-regressive prefixes from corrupting the supervisory signals during failed explorations. Extensive experiments across six benchmarks demonstrate the framework's effectiveness, achieving an average accuracy improvement of 7.4% over the base model. To the best of the authors' knowledge, this is the first work to successfully exploit on-policy self-distillation for test-time adaptation in GUI visual grounding. By filling the gap in post-deployment adaptation, the framework completes the self-evolving capability of GUI agents. The code will be released.

The paper notes that existing GUI grounding models remain static after deployment, with model parameters frozen once training is completed, and recent test-time adaptation methods rely on sparse scalar rewards that provide little information about why a prediction fails or how the model should correct it. The proposed framework equips the agent with four core capabilities: Exploration in unknown interfaces, Evaluation of generated coordinates, Reflection upon evaluation results, and Internalization of these reflections into its parameters. The Reflector takes the screenshot, instruction, and predicted coordinates as input to output an evaluation score S and a detailed reasoning process R, where S ∈ 0, 1 indicates whether the current exploration is successful and R represents the reasoning behind this evaluation. The Reflector is trained using GRPO on an offline collected dataset with format and binary rewards, and remains frozen throughout test-time adaptation.

For the internalization stage, R-OPSD constructs a self-teacher by utilizing the grounding model itself conditioned on the evaluation result S and reflection R as privileged information, computing token-level advantages as the log ratio of the teacher's probability to the student's probability for each generated coordinate token. The Contrastive Calibration method addresses the issue that during failed explorations, incorrect prefixes cause subsequent tokens to drift, making supervisory signals unreliable. It introduces an inverse-prompted student that is misled to treat the prediction as successful, producing a negative advantage that suppresses the initial error, while as errors accumulate, the conditioned incorrect prefix forces both models' distributions to align, decaying the advantage to near zero. Additionally, direction-based advantage clamping aligns token-level updates with the evaluation result, and the token-level advantage can be integrated with the query-level advantage from GRPO.

The experiments are conducted across ScreenSpot, ScreenSpot-v2, ScreenSpot-Pro, MMBench-GUI, OSWorld-G, and OSWorld-G-Refine benchmarks, using Element Accuracy as the evaluation metric. The method is built upon Qwen2.5-VL and Qwen3-VL, with both the grounding model and Reflector sharing a base model and fine-tuned using LoRA, alternating between roles by switching active LoRA adapters. Results show that when deployed on the unseen SSv2 dataset, the framework boosts the average accuracy of Qwen2.5-VL-3B from 50.2% to 57.4%, and for Qwen3-VL-2B, it raises the average accuracy to 69.4% (+3.7%) and 70.3% (+4.6%) when adapting on SSv2 and MMG respectively. The ablation study demonstrates that incorporating Reflection yields substantial performance gains, and Contrastive Calibration turns negative transfer into massive gains, reaching 64.3% on MMG. The comparison with standard RL methods shows that R-OPSD overcomes the limitation of GRPO which fails to optimize from failed groups, and the sensitivity analysis indicates that λ = 0.2 leads to the best performance. The framework also demonstrates scalability to larger base models, with Qwen2.5-VL-7B achieving accuracy improvements from 68.2% to 79.2% when adapting on MMG.

Improvements for AI systems

Improvements to AI systems:

  1. Self-evolving GUI agents – AI systems can now adapt to unseen interfaces after deployment without human labels, continuously improving their visual grounding accuracy (e.g., +7.4% average across six benchmarks) by exploring, evaluating, reflecting, and internalizing corrections in a closed loop.

  2. Reflection-guided token-level supervision – Instead of sparse scalar rewards, the system uses a frozen MLLM-based Reflector to generate detailed reasoning (e.g., the button is left of the icon), which is converted into dense token-level advantages via on-policy self-distillation. This enables fine-grained correction of each predicted coordinate token, not just a whole-prediction score.

  3. Contrastive Calibration for failed explorations – The system prevents incorrect auto-regressive prefixes from corrupting learning by using an inverse-prompted student (misled to treat failures as successes) to generate negative advantages that suppress initial errors, while decaying the advantage as errors accumulate. This turns negative transfer into positive gains (e.g., from 0% to +64.3% on MMG in ablation).

  4. Direction-based advantage clamping – Token-level updates are aligned with the evaluation result (success/failure) and can be integrated with query-level advantages from GRPO, allowing the system to optimize even from failed exploration groups, which standard RL methods cannot do.

  5. Scalable test-time adaptation – The framework works across model sizes (2B, 3B, 7B) and architectures (Qwen2.5-VL, Qwen3-VL), with LoRA-based role switching between the grounding model and Reflector, enabling efficient adaptation without full fine-tuning.

What the improved AI system can do:

  • Deploy a GUI agent that improves its own click/coordinate prediction accuracy on new, unseen apps or websites in real time, without any human-annotated data.

  • Provide interpretable reasoning for each grounding decision (e.g., I clicked the wrong element because it was partially occluded; the target is 10px to the right).

  • Recover from initial failures by learning from its own mistakes, even when the model generates a wrong sequence of tokens, via contrastive calibration.

  • Outperform static models and standard RL fine-tuning on benchmarks like ScreenSpot-Pro, OSWorld-G, and MMBench-GUI, achieving up to 79.2% accuracy on a 7B model (from 68.2% baseline).

  • Operate in a fully autonomous loop: explore new interfaces → evaluate its own outputs → reflect on why it failed → update its weights → repeat, enabling lifelong learning for digital assistants, robotic control, and UI automation.

Sources

Related papers