Right Screen, Wrong Transition: World Models as Verifiers for GUI Agents
cs.CV
Submitted: 2026-10-08
Updated: 2026-10-08
Project page: https://jiamingzhang.netlify.app/lgwm/ABSTRACT
Terminology
Sources
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Revisiting Feature Prediction for Learning Visual Representations from Video
- MobileDreamer: Generative Sketch World Model for GUI Agent
- DECEPTICON: How Dark Patterns Manipulate Web Agents
- The Value Equivalence Principle for Model-Based Reinforcement Learning
- Generative Visual Code Mobile World Models
- MobileWorldBench: Towards Semantic World Modeling For Mobile Agents
- On the Effects of Data Scale on UI Control Agents
- Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration
- Scaling GUI Agents with Visual State Transitions
- GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices
- ViMo: A Generative Visual GUI World Model for App Agents
- DINOv2: Learning Robust Visual Features without Supervision
- Mind the Gap: Action Rebinding Attacks against Android GUI Agents
- Android in the Wild: A Large-Scale Dataset for Android Device Control
- Benchmarking Web Agent Safety under E-commerce Deceptive Interfaces
- AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions
- OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows
- AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models