SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety
cs.CL, cs.AI, cs.SE
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/jc-244/SCOUT
Terminology
Sources
- Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents
- SafePred: A Predictive Guardrail for Computer-Using Agents via World Models
- ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning
- Agentic Reward Modeling: Verifying GUI Agent via Progressive Trajectory-Grounded Interaction
- The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
- Tree Search for Language Model Agents
- CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- The Art of Building Verifiers for Computer Use Agents
- Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
- Kimi K2.5: Visual Agentic Intelligence
- GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning
- Task-Adaptive Rubrics for GUI Reward Modeling
- A-MEM: Agentic Memory for LLM Agents
- ZeroGUI: Automating Online GUI Learning at Zero Human Cost
- ReAct: Synergizing Reasoning and Acting in Language Models
- ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory Learning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering