AgentLens: Adaptive Visual Modalities for Human-Agent Interaction in Mobile GUI Agents
cs.HC, cs.AI, cs.MA
Submitted: 2026-04-22
Updated: 2026-09-07
Code: https://github.com/google/A2UI
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Mobile GUI agents can automate smartphone tasks by interacting directly with app interfaces, but how they should communicate with users during execution remains underexplored.
Terminology
Abstract
Mobile GUI agents can automate smartphone tasks by interacting directly with app interfaces, but how they should communicate with users during execution remains underexplored. Existing systems rely on two extremes: foreground execution, which maximizes transparency but prevents multitasking, and background execution, which supports multitasking but provides little visual awareness. Through iterative formative studies, we found that users prefer a hybrid model with just-in-time visual interaction, but the most effective visualization modality depends on the task. Motivated by this, we present AgentLens, a mobile GUI agent that adaptively uses three visual modalities during human-agent interaction: Full UI, Partial UI, and GenUI. AgentLens extends a standard mobile agent with adaptive communication actions and uses Virtual Display to enable background execution with selective visual overlays. In a controlled study with 21 participants, AgentLens was preferred by 85.7% of participants and achieved the highest usability (1.94 Overall PSSUQ) and adoption-intent (6.43/7).
Sources
- Language Models are Few-Shot Learners
- Generative Interfaces for Language Models
- Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment
- MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents
- Segment Anything
- CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation
- Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
- Android in the Zoo: Chain-of-Action-Thought for GUI Agents
- MobileExperts: A Dynamic Tool-Enabled Agent Team in Mobile Devices
- AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- AutoDroid-V2: Boosting SLM-based GUI Agents via Code Generation
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support