SEAD: A State-Based Perspective on Attack and Defense in Tool-Using Agents
cs.CR, cs.AI
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/EverywhereSafety/SEAD
Project page: https://everywheresafety.github.io/sead/1
Terminology
Sources
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning
- AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
- AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks
- STAC: When Innocent Tools Form Dangerous Chains for LLM Agents
- Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
- AgentBench: Evaluating LLMs as Agents
- ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback
- AIRGuard: Guarding Agent Actions with Runtime Authority Control
- Identifying the Risks of LM Agents with an LM-Emulated Sandbox
- One Turn Too Late: Learning When to Intervene Against Multi-Turn Malicious Intent
- The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search
- TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking
- Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
- StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
- WebArena: A Realistic Web Environment for Building Autonomous Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs