AutoGUIWorld: Image Generators as Visual World Models for GUI Agent
cs.CV
Submitted: 2026-10-01
Updated: 2026-10-01
Code: https://github.com/ImYangC7/AutoGUIWorld
Terminology
Sources
- Gym-Anything: Turn any Software into an Agent Environment
- ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments
- Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
- MobileDreamer: Generative Sketch World Model for GUI Agent
- Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
- General Agentic Planning Through Simulative Reasoning with World Models
- Mind2Web: Towards a Generalist Agent for the Web
- DynaWeb: Model-Based Reinforcement Learning of Web Agents
- CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design
- CWM: An Open-Weights LLM for Research on Code Generation with World Models
- SEE: Structure-aware Exploring \& Exploiting for Long-horizon GUI Agent Trajectory Synthesis
- WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model
- GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments
- Image Generators are Generalist Vision Learners
- Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents
- Computer-Using World Model
- WebSTAR: Scalable Data Synthesis for Computer Use Agents with Step-Level Filtering
- Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
- TreeCUA: Efficiently Scaling GUI Automation with Tree-Structured Verifiable Evolution
- Learning with Challenges: Adaptive Difficulty-Aware Data Generation for Mobile GUI Agent Training
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models