AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents

arXiv:2608.05891 · cs.AI, cs.CL · Submitted 2026-08-06 · Read on arXiv

Weikai Xu, Yunren Feng, Haoxiang Lei, Kun Huang, Yuxuan Liu, Kang Zhao, Xiaolin Hu, Shuo Shang, Bo An

Nanyang Technological University · University of Electronic Science and Technology of China · Gaoling School of Artificial Intelligence, Renmin University of China · Xiamen University

cs.AI, cs.CL

Submitted: 2026-08-06

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 79/100

The gist: Mobile GUI agents typically operate apps through "pixel perception and touch actions, making them a promising interface for collecting and improving long-horizon mobile interaction policies."

Terminology

Summary

Mobile GUI agents typically operate apps through pixel perception and touch actions, making them a promising interface for collecting and improving long-horizon mobile interaction policies. However, the authors note that real trajectories are difficult to obtain for sensitive apps and privacy-critical operations, and existing simulated environments are costly to scale up, and GUI world models still suffer from unstable generation, limited modality coverage, and inconsistent action-transition logic.

The paper identifies three requirements that limit the usefulness of GUI world models: (1) Stable Fidelity — when the agent performs complex actions during trajectory collection, world models should be able to stably output rendered pages that are highly consistent with real devices; (2) Hybridization Modality — a text-only world model cannot directly provide visual training data. In contrast, an image-only model often fails to represent dense UI text and fine-grained layout; (3) Transition Logic Consistency — the predicted transition must follow the interaction logic in real apps, which requires world models to express a refusal when faced with invalid actions.

The paper proposes AppDeltaWorld, a transition-grounded delta code world model that predicts the next GUI as a reachable code update rather than as an unconstrained image or text description. It explicitly separates structural retrieval, multimodal rendering, and action-transition validation.

The approach "first localizes the current screen to an app-specific source state, then constrains target retrieval to action-reachable screen families, and finally retrieves a Level-1 HTML reference before generating the Level-2 HTML code. HTML is decomposed into two levels: a Level-1 HTML state captures reusable page structure and layout, while a Level-2 HTML code completes the concrete next screen."

Given a current screenshot x t, app identifier a, and mobile action m t, the current screen is converted into a structured textual state r t = φ text(x t) = [f t, d t, l t, q t, p t] where these denote the functional template, screen description, region layout, slot schema, and page type. Each historical screen is stored in a memory index as a tuple (r i, h i(1), c i(0), c i(1)). An offline index assigns each screen to a coarse functional category via a rule-based classifier c i(0) = g(p i, f i, l i, q i), then builds separate hashed TF–IDF vectors for layout, DOM signature, semantic text, and slot schema combined as: z i = 0.45·z i layout + 0.25·z i dom + 0.20·z i sem + 0.10·z i slot, with cosine-similarity clustering producing fine-grained clusters c i(1).

To preserve action logic, AppDeltaWorld constrains target retrieval with a transition index before selecting a reference HTML. For each source cluster c t and action m t, the transition memory stores candidate target clusters. Click and long-press actions quantize the action target into a 6×12 grid; swipe actions can use start and end grids. The allowed target set is C t+1 = T(c t, α(m t), κ(m t)), where α(m t) is the action type and κ(m t) is the action-target key. A semantic next-screen state ŝ t+1 is predicted, forming a target query q t+1 tar = [a, u, m t, ŝ t+1], and final Level-1 reference selection uses constrained retrieval: i* = argmax over i with a i = a and c i(1) ∈ C t+1 of ⟨e(q t+1 tar), e(r i)⟩. Transitions without supported target clusters are rejected as invalid or low-confidence.

AppDeltaWorld combines text for semantic retrieval and transition grounding, code for executable layout, and diffusion-based synthesis for visual slots. Next-state prediction is formalized as: x̂ t+1 = R(I(G θ(x t, m t, ŝ t+1, h i*(1)))), where ŝ t+1 is the predicted next-screen text, G θ is the multimodal HTML world model, I inserts generated illustration assets, and R renders Level-2 HTML into a screenshot. The code stage generates a Level-2 HTML screen as a delta from the retrieved Level-1 reference, with the generator receiving the current screenshot, the structured action record, the predicted next-screen text, and the retrieved reference HTML. The visual stage uses Qwen-Image-style text-to-image synthesis to fill image slots from textual descriptions, since on product pages and in video applications, the absence of visual assets in image slots reduces the prediction fidelity of the world model.

The authors emphasize that unlike previous work, we are more concerned about whether AppDeltaWorld can serve as a substitute for the real environment, providing learnable experience for action policy models. Rollout construction starts from real task seeds: We sample initial screenshots and instructions from GUI-Owl and OpenMobile, store them as rollout seeds, and then run an action model inside the AppDeltaWorld environment. A seed is ξ = (x 0, u, a). The action model predicts m t ∼ π η(· x t, u, h t), and AppDeltaWorld predicts and renders the next observation. The rendered PNG is then fed back to the action model for the next step, producing a closed-loop synthetic trajectory.

Quality filtering requires: the action description policy and the action coordinate must be consistent; continuous repetitive actions are not accepted; and the final step must be an active termination rather than reaching the maximum steps. The accepted trajectory condition is: D sft = (x t, u, h t, m t) parse(m t) = 1, blank(x t+1) = 0 or m t ∈ M term. Only 1/10 of the data passed the quality verification, with main failure reasons being: (1) Task Progress Loss, (2) Poor Simulation Fidelity after increased link steps, and (3) tasks Not Completed within the specified number of steps.

The world-model training data comprises 100,149 GUI transition steps: CMGUI (95,614 steps, 95.47%), CAGUI (2,978, 2.97%), Magic-RICH (1,304, 1.30%), and ChiM-Nav (253, 0.25%). All data was reverse engineered using Claude-4.8-Opus and Gemini-3.1-Pro to obtain renderable code; the code contains Level 1- and Level 2- labels to distinguish the structure and components. The action-model training data includes GUI-Owl (10% sampled, 56,237 samples, 48.18%), AppDelta trajectories constructed by AppDeltaWorld (33,133, 28.38%), and OpenMobile (27,360, 23.44%).

Under Code2World evaluation, AppDeltaWorld achieves the SoTA overall score of 73.51, outperforming all evaluated proprietary multimodal baselines, including GPT-Image-2 and Gemini-3.1-Pro-Image. Compared with the Qwen3-8B base model (50.53), this is a 22.98 absolute improvement, a 45.5% relative gain. AppDeltaWorld obtains 56.26/57.43 on S ele/S lay versus 44.00/41.70 for GPT-Image-2. Notably, GPT-Image-2 shows better element alignment than AppDeltaWorld per manual evaluation.

Ablation results: Removing diffusion decreases the overall score from 73.51 to 70.91, with the largest drops in S ele and S lay (56.26/57.43 to 50.00/50.50). Removing RAG causes a larger decline to 67.46 and substantially reduces S ad/S id from 79.69/77.00 to 65.16/69.40. Removing both yields the lowest overall score (65.68).

AndroidLens: AppDeltaAgent-8B achieves the highest AMS across all language and instruction-granularity splits. Relative to Qwen3-VL-8B, it improves Total-LL AMS/ATP from 80.33/34.96 to 90.28/46.63 (33.4% relative ATP gain) and Total-HL AMS/ATP from 78.08/23.30 to 82.53/33.05 (41.8% relative ATP gain).

MobileGym: AppDeltaAgent-8B achieves an SR of 14.1, representing a relative improvement of 38.2% over the baseline model Qwen3-VL-8B-Instruct (10.2), while increasing PR from 22.0 to 26.4. Improvements are most pronounced on L1 (66.2→83.8) and L2 (13.4→19.5) tasks.

MobileWorld: AppDeltaAgent-8B achieves a GUI-only SR of 14.9, improving Qwen3-VL-8B-Instruct (9.4) by 58.5%, while increasing the average number of interaction steps from 24.8 to 30.1.

In self-score RL, for each current-screen and instruction pair, the policy rollouts eight actions, and AppDeltaWorld predicts and renders the next state, with the policy self-scoring action consistency and instruction progress. The Step-48 policy improves AMS by 1.51, 3.25, 3.02, and 5.06 points on Feizhu, Meituan, JD, and Taobao, respectively, with mean reward rising from 68 to 76 over the first 60 updates.

In clustering-reward RL, predicted states are clustered using a 0.60 visual-similarity and 0.40 text-similarity score, with threshold 0.82, with rewards given only if the predicted state belongs to the unique largest cluster with at least two members. The consensus reward averages 0.412, but average winning-cluster support is only 3.295 of 8 rollouts and 23.5% of rollout groups have no unique supported winner on average, revealing a limitation of the current dense visual–text state representation: semantically equivalent successor states are not yet sufficiently consistent to form reliable clusters.

Action-Type Improvements: The gains are not driven by coordinate localization. On low-level instructions, the largest gains come from tool-oriented actions: wait improves from 14.19 to 72.97, type from 77.66 to 89.37, but click only from 91.82 to 95.88. On high-level instructions, swipe rises from 41.31 to 58.37, stop from 43.73 to 59.32. The authors conclude that even when page-element positions do not perfectly align with those in real apps, the higher-level interaction experience remains valuable for policy learning.

SFT Data Scaling: Public SFT data provides the foundation (improving Total-LL/HL ATP by 5.83/4.87 over base). AppDeltaWorld rollouts add complementary supervision, raising Total-LL ATP from 40.79 to 46.63 and Total-HL ATP from 28.17 to 33.05 when combined with public data. However, "AppDeltaWorld data cannot be used effectively in isolation because the world model still exhibits biases in fundamental interaction patterns; under the ADW-only setting, Total-LL/HL ATP instead decreases by 1.16/1.15 relative to the base. Performance improves monotonically from 0K to 28K added data but with diminishing returns: approximately 75% of the final gain is achieved by 12K, and performance saturates after 20K."

The paper's contributions are: (1) transition-grounded hierarchical HTML code RAG, which retrieves Level-1 reference code and uses app-specific action transitions to constrain Level-2 next-screen generation; (2) a hybrid text-code-diffusion world model that combines semantic retrieval, executable HTML generation, and image-slot rendering; and (3) validation on CMGUIBench-500 showing world-model-in-the-loop SFT and RL-training improve downstream AppDeltaAgent evaluation. The authors conclude that a transition-grounded GUI world model can serve as a scalable complement to real interaction environments for generating useful policy supervision.

Improvements for AI systems

Improvements I can make:

  • Build a transition-grounded delta-code world model that predicts the next GUI state as an executable HTML delta rather than an unconstrained image or text. The model first localizes the current screen to an app-specific source state, constrains candidate next screens to action-reachable transitions, retrieves a Level-1 HTML template, and generates Level-2 HTML for the concrete next screen. This makes the world model stable, refuses invalid actions, and preserves real app interaction logic.

  • Add a hybrid text–code–diffusion rendering pipeline. Use text for semantic retrieval, HTML for executable layout and dense UI text, and diffusion-based synthesis for image slots. This lets the system render visually faithful screenshots with accurate tiny text, structured layout, and realistic product/video assets—something text-only or image-only world models cannot do.

  • Enable world-model-in-the-loop synthetic trajectory generation. Starting from real task seeds (screenshots + instructions), run a policy inside the world model to produce closed-loop rollouts, then filter trajectories by action-parse consistency, non-repetitiveness, and active termination. This creates large-scale, learnable policy supervision without collecting sensitive real-app interactions.

  • Use the world model for RL training by rolling out multiple actions per state. I can implement self-score RL where the policy scores action consistency and instruction progress on predicted next states, yielding measurable AMS gains (e.g., +1.5 to +5.1 points on real app tasks). I can also implement clustering-reward RL using visual-text similarity to give consensus rewards, and improve the dense state representation so semantically equivalent successor states cluster more reliably.

  • Improve the action policy’s high-level interaction skills by generating synthetic data targeted at tool-oriented actions—wait, type, swipe, stop—which give the largest gains, rather than over-investing in pixel-exact click coordinates. The improved agent can learn longer-horizon behavioral patterns even when rendered page-element positions are not perfectly aligned with real devices.

  • Design a data-mix and scaling strategy for synthetic SFT data: combine public data with world-model rollouts, add roughly 12K–20K synthetic steps for most of the gain, and avoid using synthetic data alone, which otherwise introduces interaction-pattern bias. The resulting system gets monotonic improvements in task completion and action precision.

  • Decompose screens into Level-1 structural templates and Level-2 concrete components, and store them in retrievable memory with TF-IDF-plus-semantic indexing. This lets the improved system reuse page structure across transitions, retrieve references by functional category, layout, DOM, semantics, and slot schema, and generate consistent next screens with low computational cost.

  • Reverse-engineer real GUI trajectories into renderable HTML code with Level-1/Level-2 labels. This gives the improved system a scalable data pipeline for training world models on diverse apps, including sensitive domains, while retaining executable fidelity.

The improved AI system can:

  • Act as a drop-in replacement for a real mobile environment for training GUI agents, producing stable, transition-valid next-screen predictions.

  • Generate high-fidelity hybrid observations—semantically correct text, executable layouts, and realistic visual assets.

  • Reject invalid actions by refusing transitions that do not exist in real app logic.

  • Train downstream agents via SFT and RL in the synthetic world, improving task success rate, action precision, and long-horizon planning without additional real-device interaction data.

Sources

Related papers