AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents
Weikai Xu, Yunren Feng, Haoxiang Lei, Kun Huang, Yuxuan Liu, Kang Zhao, Xiaolin Hu, Shuo Shang, Bo An
Nanyang Technological University · University of Electronic Science and Technology of China · Gaoling School of Artificial Intelligence, Renmin University of China · Xiamen University
cs.AI, cs.CL
Submitted: 2026-08-06
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 79/100
The gist: Mobile GUI agents typically operate apps through "pixel perception and touch actions, making them a promising interface for collecting and improving long-horizon mobile interaction policies."
Terminology
Summary
Mobile GUI agents typically operate apps through pixel perception and touch actions, making them a promising interface for collecting and improving long-horizon mobile interaction policies.
However, the authors note that real trajectories are difficult to obtain for sensitive apps and privacy-critical operations,
and existing simulated environments are costly to scale up, and GUI world models still suffer from unstable generation, limited modality coverage, and inconsistent action-transition logic.
The paper identifies three requirements that limit the usefulness of GUI world models: (1) Stable Fidelity — when the agent performs complex actions during trajectory collection, world models should be able to stably output rendered pages that are highly consistent with real devices
; (2) Hybridization Modality — a text-only world model cannot directly provide visual training data. In contrast, an image-only model often fails to represent dense UI text and fine-grained layout
; (3) Transition Logic Consistency — the predicted transition must follow the interaction logic in real apps, which requires world models to express a refusal when faced with invalid actions.
The paper proposes AppDeltaWorld, a transition-grounded delta code world model that predicts the next GUI as a reachable code update rather than as an unconstrained image or text description.
It explicitly separates structural retrieval, multimodal rendering, and action-transition validation.
The approach "first localizes the current screen to an app-specific source state, then constrains target retrieval to action-reachable screen families, and finally retrieves a Level-1 HTML reference before generating the Level-2 HTML code. HTML is decomposed into two levels:
a Level-1 HTML state captures reusable page structure and layout, while a Level-2 HTML code completes the concrete next screen."
Given a current screenshot x t, app identifier a, and mobile action m t, the current screen is converted into a structured textual state r t = φ text(x t) = [f t, d t, l t, q t, p t]
where these denote the functional template, screen description, region layout, slot schema, and page type.
Each historical screen is stored in a memory index as a tuple (r i, h i(1), c i(0), c i(1)). An offline index assigns each screen to a coarse functional category via a rule-based classifier c i(0) = g(p i, f i, l i, q i), then builds separate hashed TF–IDF vectors for layout, DOM signature, semantic text, and slot schema
combined as: z i = 0.45·z i layout + 0.25·z i dom + 0.20·z i sem + 0.10·z i slot, with cosine-similarity clustering producing fine-grained clusters c i(1).
To preserve action logic, AppDeltaWorld constrains target retrieval with a transition index before selecting a reference HTML. For each source cluster c t and action m t, the transition memory stores candidate target clusters.
Click and long-press actions quantize the action target into a 6×12 grid; swipe actions can use start and end grids. The allowed target set is C t+1 = T(c t, α(m t), κ(m t)), where α(m t) is the action type and κ(m t) is the action-target key. A semantic next-screen state ŝ t+1 is predicted, forming a target query q t+1 tar = [a, u, m t, ŝ t+1], and final Level-1 reference selection uses constrained retrieval: i* = argmax over i with a i = a and c i(1) ∈ C t+1 of ⟨e(q t+1 tar), e(r i)⟩. Transitions without supported target clusters are rejected as invalid or low-confidence.
AppDeltaWorld combines text for semantic retrieval and transition grounding, code for executable layout, and diffusion-based synthesis for visual slots.
Next-state prediction is formalized as: x̂ t+1 = R(I(G θ(x t, m t, ŝ t+1, h i*(1)))), where ŝ t+1 is the predicted next-screen text, G θ is the multimodal HTML world model, I inserts generated illustration assets, and R renders Level-2 HTML into a screenshot. The code stage generates a Level-2 HTML screen as a delta from the retrieved Level-1 reference,
with the generator receiving the current screenshot, the structured action record, the predicted next-screen text, and the retrieved reference HTML.
The visual stage uses Qwen-Image-style text-to-image synthesis
to fill image slots from textual descriptions, since on product pages and in video applications, the absence of visual assets in image slots reduces the prediction fidelity of the world model.
The authors emphasize that unlike previous work, we are more concerned about whether AppDeltaWorld can serve as a substitute for the real environment, providing learnable experience for action policy models.
Rollout construction starts from real task seeds: We sample initial screenshots and instructions from GUI-Owl and OpenMobile, store them as rollout seeds, and then run an action model inside the AppDeltaWorld environment.
A seed is ξ = (x 0, u, a). The action model predicts m t ∼ π η(· x t, u, h t), and AppDeltaWorld predicts and renders the next observation. The rendered PNG is then fed back to the action model for the next step, producing a closed-loop synthetic trajectory.
Quality filtering requires: the action description policy and the action coordinate must be consistent; continuous repetitive actions are not accepted; and the final step must be an active termination rather than reaching the maximum steps.
The accepted trajectory condition is: D sft = (x t, u, h t, m t) parse(m t) = 1, blank(x t+1) = 0 or m t ∈ M term. Only 1/10 of the data passed the quality verification,
with main failure reasons being: (1) Task Progress Loss, (2) Poor Simulation Fidelity after increased link steps, and (3) tasks Not Completed within the specified number of steps.
The world-model training data comprises 100,149 GUI transition steps: CMGUI (95,614 steps, 95.47%), CAGUI (2,978, 2.97%), Magic-RICH (1,304, 1.30%), and ChiM-Nav (253, 0.25%). All data was reverse engineered using Claude-4.8-Opus and Gemini-3.1-Pro to obtain renderable code; the code contains Level 1- and Level 2- labels to distinguish the structure and components.
The action-model training data includes GUI-Owl (10% sampled, 56,237 samples, 48.18%), AppDelta trajectories constructed by AppDeltaWorld (33,133, 28.38%), and OpenMobile (27,360, 23.44%).
Under Code2World evaluation, AppDeltaWorld achieves the SoTA overall score of 73.51, outperforming all evaluated proprietary multimodal baselines, including GPT-Image-2 and Gemini-3.1-Pro-Image.
Compared with the Qwen3-8B base model (50.53), this is a 22.98 absolute improvement, a 45.5% relative gain. AppDeltaWorld obtains 56.26/57.43 on S ele/S lay versus 44.00/41.70 for GPT-Image-2. Notably, GPT-Image-2 shows better element alignment than AppDeltaWorld
per manual evaluation.
Ablation results: Removing diffusion decreases the overall score from 73.51 to 70.91, with the largest drops in S ele and S lay (56.26/57.43 to 50.00/50.50). Removing RAG causes a larger decline to 67.46 and substantially reduces S ad/S id from 79.69/77.00 to 65.16/69.40. Removing both yields the lowest overall score (65.68).
AndroidLens: AppDeltaAgent-8B achieves the highest AMS across all language and instruction-granularity splits.
Relative to Qwen3-VL-8B, it improves Total-LL AMS/ATP from 80.33/34.96 to 90.28/46.63 (33.4% relative ATP gain) and Total-HL AMS/ATP from 78.08/23.30 to 82.53/33.05 (41.8% relative ATP gain).
MobileGym: AppDeltaAgent-8B achieves an SR of 14.1, representing a relative improvement of 38.2% over the baseline model Qwen3-VL-8B-Instruct (10.2), while increasing PR from 22.0 to 26.4.
Improvements are most pronounced on L1 (66.2→83.8) and L2 (13.4→19.5) tasks.
MobileWorld: AppDeltaAgent-8B achieves a GUI-only SR of 14.9, improving Qwen3-VL-8B-Instruct (9.4) by 58.5%, while increasing the average number of interaction steps from 24.8 to 30.1.
In self-score RL, for each current-screen and instruction pair, the policy rollouts eight actions, and AppDeltaWorld predicts and renders the next state,
with the policy self-scoring action consistency and instruction progress. The Step-48 policy improves AMS by 1.51, 3.25, 3.02, and 5.06 points on Feizhu, Meituan, JD, and Taobao, respectively, with mean reward rising from 68 to 76 over the first 60 updates.
In clustering-reward RL, predicted states are clustered using a 0.60 visual-similarity and 0.40 text-similarity score, with threshold 0.82,
with rewards given only if the predicted state belongs to the unique largest cluster with at least two members. The consensus reward averages 0.412, but average winning-cluster support is only 3.295 of 8 rollouts
and 23.5% of rollout groups have no unique supported winner on average,
revealing a limitation of the current dense visual–text state representation: semantically equivalent successor states are not yet sufficiently consistent to form reliable clusters.
Action-Type Improvements: The gains are not driven by coordinate localization.
On low-level instructions, the largest gains come from tool-oriented actions: wait improves from 14.19 to 72.97, type from 77.66 to 89.37, but click only from 91.82 to 95.88.
On high-level instructions, swipe rises from 41.31 to 58.37, stop from 43.73 to 59.32.
The authors conclude that even when page-element positions do not perfectly align with those in real apps, the higher-level interaction experience remains valuable for policy learning.
SFT Data Scaling: Public SFT data provides the foundation (improving Total-LL/HL ATP by 5.83/4.87 over base). AppDeltaWorld rollouts add complementary supervision, raising Total-LL ATP from 40.79 to 46.63 and Total-HL ATP from 28.17 to 33.05 when combined with public data. However, "AppDeltaWorld data cannot be used effectively in isolation because the world model still exhibits biases in fundamental interaction patterns; under the ADW-only setting, Total-LL/HL ATP instead decreases by 1.16/1.15 relative to the base. Performance improves monotonically from 0K to 28K added data but with diminishing returns:
approximately 75% of the final gain is achieved by 12K, and performance saturates after 20K."
The paper's contributions are: (1) transition-grounded hierarchical HTML code RAG, which retrieves Level-1 reference code and uses app-specific action transitions to constrain Level-2 next-screen generation
; (2) a hybrid text-code-diffusion world model that combines semantic retrieval, executable HTML generation, and image-slot rendering
; and (3) validation on CMGUIBench-500 showing world-model-in-the-loop SFT and RL-training improve downstream AppDeltaAgent evaluation.
The authors conclude that a transition-grounded GUI world model can serve as a scalable complement to real interaction environments for generating useful policy supervision.
Improvements for AI systems
Improvements I can make:
-
Build a transition-grounded delta-code world model that predicts the next GUI state as an executable HTML delta rather than an unconstrained image or text. The model first localizes the current screen to an app-specific source state, constrains candidate next screens to action-reachable transitions, retrieves a Level-1 HTML template, and generates Level-2 HTML for the concrete next screen. This makes the world model stable, refuses invalid actions, and preserves real app interaction logic.
-
Add a hybrid text–code–diffusion rendering pipeline. Use text for semantic retrieval, HTML for executable layout and dense UI text, and diffusion-based synthesis for image slots. This lets the system render visually faithful screenshots with accurate tiny text, structured layout, and realistic product/video assets—something text-only or image-only world models cannot do.
-
Enable world-model-in-the-loop synthetic trajectory generation. Starting from real task seeds (screenshots + instructions), run a policy inside the world model to produce closed-loop rollouts, then filter trajectories by action-parse consistency, non-repetitiveness, and active termination. This creates large-scale, learnable policy supervision without collecting sensitive real-app interactions.
-
Use the world model for RL training by rolling out multiple actions per state. I can implement self-score RL where the policy scores action consistency and instruction progress on predicted next states, yielding measurable AMS gains (e.g., +1.5 to +5.1 points on real app tasks). I can also implement clustering-reward RL using visual-text similarity to give consensus rewards, and improve the dense state representation so semantically equivalent successor states cluster more reliably.
-
Improve the action policy’s high-level interaction skills by generating synthetic data targeted at tool-oriented actions—
wait,type,swipe,stop—which give the largest gains, rather than over-investing in pixel-exact click coordinates. The improved agent can learn longer-horizon behavioral patterns even when rendered page-element positions are not perfectly aligned with real devices. -
Design a data-mix and scaling strategy for synthetic SFT data: combine public data with world-model rollouts, add roughly 12K–20K synthetic steps for most of the gain, and avoid using synthetic data alone, which otherwise introduces interaction-pattern bias. The resulting system gets monotonic improvements in task completion and action precision.
-
Decompose screens into Level-1 structural templates and Level-2 concrete components, and store them in retrievable memory with TF-IDF-plus-semantic indexing. This lets the improved system reuse page structure across transitions, retrieve references by functional category, layout, DOM, semantics, and slot schema, and generate consistent next screens with low computational cost.
-
Reverse-engineer real GUI trajectories into renderable HTML code with Level-1/Level-2 labels. This gives the improved system a scalable data pipeline for training world models on diverse apps, including sensitive domains, while retaining executable fidelity.
The improved AI system can:
-
Act as a drop-in replacement for a real mobile environment for training GUI agents, producing stable, transition-valid next-screen predictions.
-
Generate high-fidelity hybrid observations—semantically correct text, executable layouts, and realistic visual assets.
-
Reject invalid actions by refusing transitions that do not exist in real app logic.
-
Train downstream agents via SFT and RL in the synthetic world, improving task success rate, action precision, and long-horizon planning without additional real-device interaction data.
Sources
- MobileDreamer: Generative Sketch World Model for GUI Agent
- Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
- STEP: Success-Rate-Aware Trajectory-Efficient Policy Optimization
- OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis
- Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
- Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents
- UI-Venus Technical Report: Building High-performance UI Agents with RFT
- Computer-Using World Model
- CogAgent: A Visual Language Model for GUI Agents
- MobileIPL: Enhancing Mobile Agents Thinking Process via Iterative Preference Learning
- R^3: Replay, Reflection, and Ranking Rewards for LLM Reinforcement Learning
- Generative Visual Code Mobile World Models
- MobileWorldBench: Towards Semantic World Modeling For Mobile Agents
- From Word to World: Can Large Language Models be Implicit Text-based World Models?
- AppAgent v2: Advanced Agent for Flexible Mobile Interactions
- CoME: Empowering Channel-of-Mobile-Experts with Informative Hybrid-Capabilities Reasoning
- ViMo: A Generative Visual GUI World Model for App Agents
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- Falcon-UI: Understanding GUI Before Following User Instructions
- MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection