CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications

arXiv:2608.11588 · cs.AI · Submitted 2026-08-12 · Read on arXiv

Linqiang Guo, Li Gu, Zihuan Jiang, Zhixiang Chi, Siobhan Reid, Ziqiang Wang, Yuanhao Yu, Wei Liu, Yang Wang, Tse-Hsun, Chen

Concordia University · Mila – Québec AI Institute · University of Toronto · McMaster University

cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 50/100

The gist: Mobile GUI agents "remain brittle when deployed to applications absent from source training." The paper studies "novel-app generalization under a limited target interaction budget and without target

Terminology

Summary

Mobile GUI agents remain brittle when deployed to applications absent from source training. The paper studies novel-app generalization under a limited target interaction budget and without target demonstrations. The authors note that an agent may instead encounter an application whose interface and task procedures were absent from training, and adapting to such applications from limited interaction is therefore essential for effective operation beyond the original training environment.

The paper highlights a key gap: "An unseen application can expose both interface-specific and procedural gaps. An agent may understand the goal but fail to ground actions to unfamiliar interface elements, or it may execute individual actions correctly yet lack the workflow and completion conditions needed to finish the task. Existing work only partially addresses this target-side adaptation problem — AndroidWorld-Generalization updates the policy from the agent's own target-app rollouts while leaving its workflow context fixed, and UI-Mem jointly learns experience memory and policy during source-side training, then transfers the resulting knowledge to unseen applications."

The paper introduces CoAdapt-GUI, a test-time adaptation (TTA) framework that maintains and jointly updates these two states from target-app rollouts and task-level rewards. The framework has two channels:

  1. Context channel: contrasts successful and failed traces to revise transferable procedures, failure patterns, and completion checks while excluding app-bound details.

  2. Policy channel: updates a lightweight LoRA adapter through a task–context-conditioned group-relative objective, while keeping the VLM backbone frozen.

Both updates are derived from the same rollout groups within each adaptation round, and the resulting context and policy are frozen before evaluation on held-out target tasks.

The paper separates source trajectories into two components: an app-bound state M app a and a transferable state M tr a. The app-bound state is instantiated as a screen-transition FSM recording concrete screens, action-conditioned transitions, visible interface cues, and resource-level information. The transferable state contains workflow entries of the form w = ⟨c, P, F, V⟩ where c specifies when the workflow applies, P describes an abstract procedure, F records failure or recovery conditions, and V specifies observable or executable completion checks.

An "eligibility predicate, Eligible tr(w) ∈ 0, 1, determines which entries may cross application boundaries. A schema validator and linter reject entries containing app names, package or resource identifiers, concrete widget labels, coordinates, task-instance values, and other app-bound content."

The two channels reuse the same rollout stream but operate at different frequencies. The updates are "interaction-coupled rather than jointly differentiable: the current context shapes the trajectories used for policy learning, while the current policy determines the successful and failed behaviors available for future context revisions."

For the context update, the controller maintains a population of TrueSkill-rated context variants and samples already materialized variants, evaluates them on matched tasks and resets seeds, and updates their ratings using the resulting task rewards. A frozen reflector then contrasts successful and failed traces from the evaluated variants and proposes a typed workflow revision to a high-rated parent.

For the policy update, trajectories are partitioned into groups sharing both task and context: G(q, κ) = j q j = q, κ j = κ. The advantage is computed as A j = (r j − r̄ G) / s G where s G is the normalization factor: it is set to one for mean-centered advantages, or to the within-group standard deviation plus a small constant for standardized advantages. The objective is L policy = −(1/B act) Σ j∈B act A j l j(θ) + βR(θ; π anchor) where l j(θ) averages log-probabilities over its action-generation units.

Following the released unseen-app setting, CoAdapt-GUI reaches 45.0%, exceeding the reported Policy-Only TTA baseline of 37.5% by 7.5 percentage points. Context-Only TTA reaches 35.00% ± 1.74, Static Context Transfer reaches 28.75% ± 2.28, and the Base Policy achieves 27.10%. The paper notes CoAdapt-GUI outperforms Context-Only TTA by 10.00 points, indicating that co-adapting workflow context and the policy is more effective than adapting context alone in this setting.

The paper constructs AndroidWorld Plus by extending AndroidWorld with three apps from B-MoCA and three from AndroidLab, resulting in 25 apps and 191 task templates with 12 apps with 96 templates to the source set and the remaining 13 apps with 95 templates to the disjoint target set.

CoAdapt-GUI raises overall success from 38.6% to 52.9%, a gain of 14.3 points. The cumulative context path reaches 48.1%, 9.5 points above the Base Policy, whereas Policy-Only TTA reaches 40.0%, a gain of only 1.4 points. CoAdapt-GUI is a further 4.8 points above Context-Only TTA.

On Category-Shared Apps, CoAdapt-GUI reaches 70.4%, compared with 53.7% for Policy-Only TTA and 63.9% for Context-Only TTA. On Category-Novel Apps, Static Context Transfer exactly matches the Base Policy at 29.4%. Policy-Only TTA falls to 25.5%, whereas Context-Only TTA and CoAdapt-GUI improve success to 31.4% and 34.3%.

The paper makes three contributions:

  1. Joint target-side adaptation from autonomous interaction — a TTA framework that updates separate workflow-context and policy states using only the agent's own target-app rollouts and task rewards, without target demonstrations or access to held-out evaluation signals.

  2. Transfer-constrained context–policy adaptation — separating transferable workflow knowledge from app-bound source state and coordinating two reward-guided updates from the same target interactions.

  3. Evaluation across instance- and task-type generalization — reaching 45.0% and 52.9% in the two settings, respectively, achieving the best overall result in both settings.

The paper acknowledges several limitations: Source workflows may contain interface-specific assumptions or errors introduced during reflection, causing negative transfer to a new application. A limited target-interaction budget can produce noisy or uniform rewards, making it difficult to distinguish useful context revisions and assign policy credit. Exploration on a new application may expose private information, trigger irreversible actions, or accumulate unstable updates over time. The paper addresses these through eligibility checks, matched task-reset evaluation, on-policy buffer clearing, LoRA-only updates on a frozen backbone, and conducting adaptation in resettable emulators with a fixed budget.

Improvements for AI systems

Improvements to AI Systems:

  1. Dual-channel test-time adaptation: Implement a system that maintains and jointly updates two separate knowledge states during deployment—a transferable workflow context (procedures, failure patterns, completion checks) and a lightweight policy adapter (LoRA)—rather than adapting only one. This enables the AI to simultaneously fix both what to do (procedural gaps) and how to ground actions (interface-specific gaps) in unfamiliar environments.

  2. Transfer-constrained knowledge filtering: Add an eligibility predicate with schema validation and linting that automatically rejects app-bound content (package names, resource IDs, concrete widget labels, coordinates) from transferable workflow entries. This prevents negative transfer of interface-specific assumptions when applying learned procedures to new applications.

  3. Population-based context evolution with TrueSkill ratings: Maintain a population of context variants rated via TrueSkill, sample and evaluate them on matched tasks with reset seeds, then use a frozen reflector to contrast successful and failed traces and propose typed workflow revisions to high-rated parents. This allows the AI to explore multiple procedural hypotheses and converge on robust, transferable workflows.

  4. Group-relative policy optimization with task-context conditioning: Partition trajectories into groups sharing both task and context, compute advantages normalized within each group (mean-centered or standardized), and optimize a LoRA adapter using a group-relative objective with an anchor regularization term. This provides stable policy updates even with noisy or uniform rewards from limited interactions.

  5. Interaction-coupled but separately frequent updates: Decouple the update frequencies of context and policy channels while reusing the same rollout stream. The context shapes future trajectories for policy learning, while the current policy determines which behaviors are available for context revision—creating a feedback loop without requiring joint differentiability.

  6. Frozen backbone with LoRA-only adaptation: Keep the VLM backbone frozen and update only a lightweight LoRA adapter, preventing catastrophic forgetting and unstable updates during test-time adaptation while enabling efficient fine-tuning with limited target interactions.

  7. Resettable emulator-based adaptation with fixed budgets: Conduct adaptation in resettable emulators with a fixed interaction budget, clearing on-policy buffers between rounds. This allows safe exploration of unfamiliar applications, mitigates risks of irreversible actions or private data exposure, and enables controlled evaluation on held-out tasks.

What the improved AI system can do:

  • Deploy to completely unseen applications and achieve task success without any target demonstrations, improving from 27.1% to 45.0% on new task instances and from 38.6% to 52.9% on new task templates.

  • Distinguish and separately repair procedural knowledge (workflow steps, completion conditions) from action-grounded knowledge (how to click, type, navigate) when encountering novel interfaces.

  • Adapt effectively even when only 1–2 interactions per task are available, using group-relative credit assignment to handle noisy rewards.

  • Avoid negative transfer by automatically filtering out app-specific details from learned procedures, ensuring only universally applicable knowledge crosses application boundaries.

  • Maintain stable performance on category-novel apps (34.3% success) where policy-only adaptation fails (25.5%), demonstrating robustness to both interface and procedural novelty.

  • Reuse a single rollout stream for both context revision and policy updates, making the adaptation process sample-efficient and practical for real-world deployment with limited interaction budgets.

Abstract

Mobile GUI agents remain brittle when deployed to applications absent from source training. We study novel-app generalization under a limited target interaction budget and without target demonstrations. We introduce CoAdapt-GUI, a test-time adaptation (TTA) framework that jointly adapts structured workflow context and policy from the agent's own target-app rollouts and rewards. The workflow context retains transferable procedures, failure modes, and verification rules while excluding app-bound source details. This separation allows reusable workflow knowledge to guide adaptation without transferring source-interface state. For policy adaptation, task-context-matched group-relative optimization updates a LoRA adapter on a frozen vision-language model. Across two unseen-app evaluations, CoAdapt-GUI reaches 45.0% on AndroidWorld-Generalization, compared with 37.5% for the reported Policy-Only TTA baseline, and raises AndroidWorld Plus performance from 38.6% to 52.9%. These results show that transfer-constrained workflow context provides substantial gains and that joint policy adaptation further improves held-out performance.

Sources

Related papers