Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation
Shiyu Xuan, Zechao Li
Nanjing University of Science and Technology
cs.CV, cs.AI, cs.CL
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This paper introduces a Test-Time Self-Evolving framework for GUI visual grounding that enables models to improve after deployment without human-annotated ground truth.
Terminology
Summary
This paper introduces a Test-Time Self-Evolving framework for GUI visual grounding that enables models to improve after deployment without human-annotated ground truth. The framework constructs a closed-loop of Exploration, Evaluation, Reflection, and Internalization. Specifically, the agent first explores unseen interfaces by predicting grounding coordinates for given instructions. To evaluate these explorations, the authors introduce an MLLM-based Reflector to assess the generated results and provide corresponding reasoning reflections. To internalize reflection knowledge into the model weights, they propose Reflection-Guided On-Policy Self-Distillation (R-OPSD), which translates high-level reasoning into dense token-level supervision via a conditioned self-teacher. Furthermore, they design a Contrastive Calibration method to prevent incorrect auto-regressive prefixes from corrupting the supervisory signals during failed explorations. Extensive experiments across six benchmarks demonstrate the framework's effectiveness, achieving an average accuracy improvement of 7.4% over the base model. To the best of the authors' knowledge, this is the first work to successfully exploit on-policy self-distillation for test-time adaptation in GUI visual grounding. By filling the gap in post-deployment adaptation, the framework completes the self-evolving capability of GUI agents. The code will be released.
The paper notes that existing GUI grounding models remain static after deployment, with model parameters frozen once training is completed, and recent test-time adaptation methods rely on sparse scalar rewards that provide little information about why a prediction fails or how the model should correct it. The proposed framework equips the agent with four core capabilities: Exploration in unknown interfaces, Evaluation of generated coordinates, Reflection upon evaluation results, and Internalization of these reflections into its parameters. The Reflector takes the screenshot, instruction, and predicted coordinates as input to output an evaluation score S and a detailed reasoning process R, where S ∈ 0, 1 indicates whether the current exploration is successful and R represents the reasoning behind this evaluation. The Reflector is trained using GRPO on an offline collected dataset with format and binary rewards, and remains frozen throughout test-time adaptation.
For the internalization stage, R-OPSD constructs a self-teacher by utilizing the grounding model itself conditioned on the evaluation result S and reflection R as privileged information, computing token-level advantages as the log ratio of the teacher's probability to the student's probability for each generated coordinate token. The Contrastive Calibration method addresses the issue that during failed explorations, incorrect prefixes cause subsequent tokens to drift, making supervisory signals unreliable. It introduces an inverse-prompted student that is misled to treat the prediction as successful, producing a negative advantage that suppresses the initial error, while as errors accumulate, the conditioned incorrect prefix forces both models' distributions to align, decaying the advantage to near zero. Additionally, direction-based advantage clamping aligns token-level updates with the evaluation result, and the token-level advantage can be integrated with the query-level advantage from GRPO.
The experiments are conducted across ScreenSpot, ScreenSpot-v2, ScreenSpot-Pro, MMBench-GUI, OSWorld-G, and OSWorld-G-Refine benchmarks, using Element Accuracy as the evaluation metric. The method is built upon Qwen2.5-VL and Qwen3-VL, with both the grounding model and Reflector sharing a base model and fine-tuned using LoRA, alternating between roles by switching active LoRA adapters. Results show that when deployed on the unseen SSv2 dataset, the framework boosts the average accuracy of Qwen2.5-VL-3B from 50.2% to 57.4%, and for Qwen3-VL-2B, it raises the average accuracy to 69.4% (+3.7%) and 70.3% (+4.6%) when adapting on SSv2 and MMG respectively. The ablation study demonstrates that incorporating Reflection yields substantial performance gains, and Contrastive Calibration turns negative transfer into massive gains, reaching 64.3% on MMG. The comparison with standard RL methods shows that R-OPSD overcomes the limitation of GRPO which fails to optimize from failed groups, and the sensitivity analysis indicates that λ = 0.2 leads to the best performance. The framework also demonstrates scalability to larger base models, with Qwen2.5-VL-7B achieving accuracy improvements from 68.2% to 79.2% when adapting on MMG.
Improvements for AI systems
Improvements to AI systems:
-
Self-evolving GUI agents – AI systems can now adapt to unseen interfaces after deployment without human labels, continuously improving their visual grounding accuracy (e.g., +7.4% average across six benchmarks) by exploring, evaluating, reflecting, and internalizing corrections in a closed loop.
-
Reflection-guided token-level supervision – Instead of sparse scalar rewards, the system uses a frozen MLLM-based Reflector to generate detailed reasoning (e.g.,
the button is left of the icon
), which is converted into dense token-level advantages via on-policy self-distillation. This enables fine-grained correction of each predicted coordinate token, not just a whole-prediction score. -
Contrastive Calibration for failed explorations – The system prevents incorrect auto-regressive prefixes from corrupting learning by using an inverse-prompted student (misled to treat failures as successes) to generate negative advantages that suppress initial errors, while decaying the advantage as errors accumulate. This turns negative transfer into positive gains (e.g., from 0% to +64.3% on MMG in ablation).
-
Direction-based advantage clamping – Token-level updates are aligned with the evaluation result (success/failure) and can be integrated with query-level advantages from GRPO, allowing the system to optimize even from failed exploration groups, which standard RL methods cannot do.
-
Scalable test-time adaptation – The framework works across model sizes (2B, 3B, 7B) and architectures (Qwen2.5-VL, Qwen3-VL), with LoRA-based role switching between the grounding model and Reflector, enabling efficient adaptation without full fine-tuning.
What the improved AI system can do:
-
Deploy a GUI agent that improves its own click/coordinate prediction accuracy on new, unseen apps or websites in real time, without any human-annotated data.
-
Provide interpretable reasoning for each grounding decision (e.g.,
I clicked the wrong element because it was partially occluded; the target is 10px to the right
). -
Recover from initial failures by learning from its own mistakes, even when the model generates a wrong sequence of tokens, via contrastive calibration.
-
Outperform static models and standard RL fine-tuning on benchmarks like ScreenSpot-Pro, OSWorld-G, and MMBench-GUI, achieving up to 79.2% accuracy on a 7B model (from 68.2% baseline).
-
Operate in a fully autonomous loop: explore new interfaces → evaluate its own outputs → reflect on why it failed → update its weights → repeat, enabling lifelong learning for digital assistants, robotic control, and UI automation.
Sources
- Qwen3-VL Technical Report
- Grounding Computer Use Agents on Human Demonstrations
- Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
- Entropy-Aware On-Policy Distillation of Language Models
- Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
- InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners
- Zoom to Essence: Trainless GUI Grounding by Inferring upon Interface Elements
- GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
- FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
- RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
- GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
- Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
- Self-Distilled RLVR
- Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
- Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning
- OPSDL: On-Policy Self-Distillation for Long-Context Language Models
- Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models