Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents
cs.AI
Submitted: 2026-09-03
Updated: 2026-09-03
Code: https://github.com/serein356/ConflictGuard
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Understanding intermediate layers using linear classifier probes
- Refusal in Language Models Is Mediated by a Single Direction
- Qwen3-VL Technical Report
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Finding Neurons in a Haystack: Case Studies with Sparse Probing
- Faithful Mobile GUI Agents with Guided Advantage Estimator
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- Kimi K2.5: Visual Agentic Intelligence
- UI-Venus-1.5 Technical Report
- Steering Language Models With Activation Engineering
- Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
- The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
- Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
- VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
- Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
- VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection