Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments
cs.AI
Submitted: 2026-08-25
Updated: 2026-08-25
Code: https://github.com/gguogan/AnTrap
Terminology
Sources
- LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios
- VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics
- When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
- MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented Environments
- It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents
- PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- Android in the Wild: A Large-Scale Dataset for Android Device Control
- AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the Wild
- Qwen3 Technical Report
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- OpenComputer: Verifiable Software Worlds for Computer-Use Agents
- Step-level Optimization for Efficient Computer-use Agents
- GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior
- BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism
- MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World Environment
- OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents
- Mobile-Bench-v2: A More Realistic and Comprehensive Benchmark for VLM-based Mobile Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection