SIR: Self-improving Red-teaming for Compute Use Agents
cs.CR, cs.AI
Submitted: 2026-08-31
Updated: 2026-10-03
Terminology
Sources
- Jailbreaking Black Box Large Language Models in Twenty Queries
- How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
- WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
- EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
- RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments
- AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
- EVA: Evolving Semantic Adversaries for Red-Teaming GUI Agents Against Environmental Injection Attacks
- Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
- Ignore Previous Prompt: Attack Techniques For Language Models
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
- A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions
- MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
- On Adaptive Attacks to Adversarial Example Defenses
- OpenCUA: Open Foundations for Computer-Use Agents
- AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
- Dissecting Adversarial Robustness of Multimodal LM Agents
- CoP: Agentic Red-teaming for Large Language Models using Composition of Principles
- AdvAgent: Controllable Blackbox Red-teaming on Web Agents
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs