SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation
Zhou Liu, Ligang Huang, Zeli Su, Zewei Pan, Zhaoyang Han, Xing Chen, Yuanfeng Song, Wentao Zhang
Peking University · ByteDance · Shanghai Jiao Tong University
cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 100/100
The gist: SkillLens converts heterogeneous GUI experience into Visual Skill Cards (VSCs) that provide visual procedural memory to frozen VLM executors.
Terminology
Summary
SkillLens converts heterogeneous GUI experience into Visual Skill Cards (VSCs) that provide visual procedural memory to frozen VLM executors. A VSC is a state-conditioned memory object that binds reusable procedures with applicability cues, visual evidence, and verification signals. SkillLens constructs VSCs from heterogeneous interaction experience through Trace-to-Visual-Skill-Card and, at inference time, retrieves relevant cards and selectively expands only the evidence needed by a fixed visual-language model executor for grounded GUI action prediction. The same representation also supports CardDistill, which uses VSC evidence as privileged teacher context to train a student that acts without runtime card retrieval.
The paper introduces three contributions: (1) formulating VSCs as a state-conditioned interface for visual procedural memory and introducing Trace-to-VSC to convert heterogeneous interaction sources into this common representation; (2) developing SkillLens, which decouples low-cost retrieval from selective high-resolution evidence expansion while preserving the live screen as the final grounding source; (3) introducing CardDistill, a VSC-guided on-policy distillation algorithm that transfers card-conditioned teacher behavior into a student policy without runtime card retrieval.
The VSC representation separates three roles: procedure text describes the intended operation, state cues describe when the operation applies, and visual evidence anchors the procedure to interface appearance. VSCs can be organized hierarchically when the source trace supports it: a meta card captures a long-horizon task pattern, a core card captures a reusable subgoal, and an execution card captures a short grounded operation. The reported experiments rank execution-level cards directly.
Trace-to-VSC construction follows a common flow: a source record is first normalized into a trace, the trace is segmented into reusable units, each unit is summarized into a procedure, visual evidence is bound to the procedure, and the final card is audited. The audit checks that the card is self-contained and that held-out evaluations do not expose answer coordinates as templates. Different sources instantiate the flow with different adapters, but they all export the same VSC schema.
SkillLens retrieval separates retrieval from evidence expansion. The default path uses a lightweight context-aware selector that matches the current task and interface state against the card library, then expands only the selected evidence. The retrieval query is qt = Q(x, ht, ct), where Q deterministically serializes the task instruction x, recent interaction history ht, and selector-visible interface metadata ct. For each card si, di concatenates its identifier, name, goal, keywords, and available metadata. Coarse retrieval produces a candidate set, then reranking selects the top candidates, and finally a deterministic expansion function loads the bounded textual and visual evidence of selected cards. This separation bounds runtime evidence while preserving fine-grained visual detail.
Expanded evidence is reference material, not a coordinate template. SkillLens constructs the executor context from the live screenshot, the task and recent history, the selected VSC evidence, and the action schema of the benchmark. The frozen executor then predicts at = πθ0(ξt, St, Et), where θ0 is fixed. The live observation remains the source of grounding: retrieved evidence can identify a familiar state, highlight a relevant control pattern, or provide a completion cue, but the final action must still be taken on the current screen. Verification cues define observable postconditions for a successful step.
CardDistill turns VSCs from runtime memory into training-time supervision. Each training state ξt = (x, ot, ht) is paired with source-aligned VSC evidence, represented as a bounded privileged bundle (Stpriv, Etpriv). The student policy πθ sees only the benchmark-native context ξt, while the teacher policy πϕ receives the card-conditioned context ξ˜t = (ξt, Stpriv, Etpriv). The student generates an on-policy action sequence, and the teacher evaluates the same student-generated prefixes. The reported runs optimize the teacher-confidence-weighted reverse KL loss. At evaluation time, the teacher and VSC evidence are removed, and the student predicts from ξt alone.
The paper evaluates SkillLens and CardDistill on Multimodal-Mind2Web, WebLINX-BrowserGym (WebLINX-BG), and OSWorld-G. The main results show that VSCs improve frozen VLM executors on both web-action and GUI-grounding tasks. On GPT-5.4-mini, SkillLens improves Mind2Web Step SR from 77.2 to 88.8 (+11.6), WebLINX-BG Overall from 12.8 to 15.7 (+2.9), and OSWorld-G Grounding Acc. from 45.0 to 66.5 (+21.5). GPT-4o and the Gemini models show the same positive trend across the three benchmark groups. The open Qwen3-VL-2B executor likewise improves across all three groups: Mind2Web Step SR rises from 4.6 to 10.3, WebLINX-BG Overall rises from 14.1 to 16.5 (+2.4), and OSWorld-G Grounding Acc. rises from 28.5 to 41.0.
CardDistill improves Mind2Web Step SR by 12.0 points (95% CI [7.5, 17.5]) and WebLINX-BG Overall by 3.2 points (95% CI [0.88, 5.85]). Plain OPD uses the same training setup without VSC evidence in the teacher context, while the shuffled-VSC control retains the pipeline but misaligns the cards. Neither control differs significantly from the base student in the 200-step evaluation. On Mind2Web, CardDistill outperforms shuffled VSC by 10.5 points (95% CI [5.5, 15.5], p <.001), providing evidence that source-aligned VSC evidence contributes under the matched training budget.
Selector and modality ablations show that the text-indexed selector gives the best tested trade-off in WebLINX-BG execution (17.45 Overall, 54.50 Dialog Acc., 1.5s), while VLM-board improves Element IoU from 16.96 to 18.98 but lowers Overall to 13.58. Without the executor, VLM-board improves Hit@1 on Mind2Web (0.86 to 0.93) and WebLINX-BG (0.40 to 0.55), while OCR+Text board is strong on OSWorld-G but weak on WebLINX-BG. Modality contributions are task- and executor-dependent: procedural text and visual evidence support different action-prediction and grounding demands, while the unified VSC representation provides a common interface across these settings.
Negative controls test whether the gains come from relevant VSC evidence rather than from longer prompts or extra images. The executor, prompt template, and expansion budget stay fixed, but the retrieved card is replaced with either a uniformly sampled card or a low-overlap irrelevant card. Retrieved VSCs remain much stronger: on the fixed Mind2Web subset, exact Step Acc. is 92.0 with retrieved VSCs versus 72.0 with random cards and 71.0 with irrelevant cards; on OSWorld-G, Grounding Acc. is 66.5 versus 10.5 and 12.5. This negative control shows that relevance is essential: unrelated visual evidence does not reproduce the SkillLens gain and can actively harm grounding.
A diagnostic grounding reference using a center-in-target-box score converted from native predictions shows that CardDistill falls between GUI-Actor and UI-TARS on both evaluated benchmarks under this readout. The paper concludes that relevant VSCs improve grounded action prediction, while CardDistill transfers card-conditioned behavior into a student without runtime retrieval. VSCs thus connect retrieval-augmented prediction with on-policy distillation.
Improvements for AI systems
Improvements to AI Systems:
-
Add a visual procedural memory module to frozen VLM executors. The system stores reusable, state-conditioned Visual Skill Cards (VSCs) that bind procedure text, applicability cues, and visual evidence. At inference, it retrieves only relevant cards and expands their bounded evidence into the executor’s context, improving action prediction without fine-tuning the base model.
-
Decouple retrieval from evidence expansion. A lightweight text-indexed selector matches the current task and interface state against the card library, then a deterministic expansion function loads only the selected cards’ textual and visual evidence. This keeps runtime cost low (1.5s in WebLINX-BG) while preserving high-resolution visual detail for grounding.
-
Use live screen as the final grounding source. Retrieved VSC evidence serves as reference material (identifying familiar states, highlighting control patterns, providing completion cues), but the executor always predicts actions on the current screenshot. This prevents coordinate leakage and ensures actions are valid on the actual interface.
-
Enable on-policy distillation without runtime retrieval. CardDistill uses VSC evidence as privileged teacher context during training. The teacher evaluates student-generated action prefixes, and the student learns from teacher-confidence-weighted reverse KL loss. At deployment, the student acts from the native context alone, removing retrieval overhead entirely.
-
Construct hierarchical memory for long-horizon tasks. When source traces support it, VSCs are organized as meta cards (task patterns), core cards (reusable subgoals), and execution cards (short grounded operations). This lets the system reuse subgoal procedures across different tasks and retrieve at multiple abstraction levels.
-
Add verification signals to memory. Each VSC includes observable postconditions for a successful step. The executor can check these cues on the live screen to confirm whether a step completed, improving step success rate and reducing cascading errors.
What the improved AI system can do:
-
Boost frozen VLM performance on web and GUI tasks: Step success rate on Mind2Web rises from 77.2 to 88.8 (+11.6) with GPT-5.4-mini; grounding accuracy on OSWorld-G rises from 45.0 to 66.5 (+21.5). Open Qwen3-VL-2B also improves across all benchmarks (e.g., Mind2Web Step SR from 4.6 to 10.3).
-
Train compact student policies that match teacher performance without retrieval: CardDistill improves Mind2Web Step SR by 12.0 points and WebLINX-BG Overall by 3.2 points over the base student, with no runtime card lookup.
-
Maintain robustness against irrelevant memory: Retrieved VSCs outperform random or irrelevant cards by large margins (92.0 vs 72.0 Step Acc. on Mind2Web; 66.5 vs 10.5 Grounding Acc. on OSWorld-G), showing the system only benefits from source-aligned, relevant evidence.
-
Adapt to heterogeneous interaction sources: The same VSC schema handles traces from web logs, browser gyms, and OS environments, converting them into a common memory format usable by any frozen executor.
-
Provide a unified interface for both retrieval-augmented prediction and distillation: The same VSC representation serves as runtime memory for frozen models and as training-time supervision for student policies, enabling a single pipeline for both deployment modes.
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- Thinking Without Images: Internalizing Visual Manipulation with On-Policy Self-Distillation
- CUA-Skill: Develop Skills for Computer Using Agent
- ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents
- VISUALSKILL: Multimodal Skills for Computer-Use Agents
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- DeepSeek-OCR: Contexts Optical Compression
- LensVLM: Selective Context Expansion for Compressed Visual Representation of Text
- Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills
- Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
- MMSkills: Towards Multimodal Skills for General Visual Agents
- Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection