RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems
cs.AI
Submitted: 2026-06-22
Updated: 2026-08-30
Comments: EMNLP 2026 Findings
Code: https://github.com/Azure/PyRIT
Project page: https://ibm.github.io/ares
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Agentic AI systems powered by large language models (LLMs) are rapidly evolving into autonomous decision-making systems, exposing attack vectors beyond those of traditional LLM vulnerabilities.
Terminology
Abstract
Agentic AI systems powered by large language models (LLMs) are rapidly evolving into autonomous decision-making systems, exposing attack vectors beyond those of traditional LLM vulnerabilities. Existing security evaluations are often tied to specific implementations or domains, limiting unified comparison across heterogeneous systems. To address this gap, we introduce RIFT-Bench, a representation-driven methodology for dynamic red-teaming that enables unified evaluations across diverse agentic architectures. Building on a novel hierarchical representation, RIFT-Bench operates in two automated phases: Discovery, which extracts system structure, and Scanning, which executes adaptive adversarial attacks. It directly evaluates the examined system using 105 adaptive adversarial probes spanning diverse attack vectors and objectives. We demonstrate the effectiveness of the proposed evaluation pipeline across 45 agentic systems spanning a diverse range of implementations, showing that the approach generalizes effectively to heterogeneous agentic architectures. Beyond systems and attacks, RIFT-Bench also supports direct evaluation of mitigation strategies. These key capabilities make RIFT-Bench a scalable foundation for security evaluation of agentic AI systems in practice. Infrastructure code and benchmark artifacts are available at https://tinyurl.com/RIFTBench.
Sources
- Open Agent Specification (Agent Spec): A Unified Representation for AI Agents
- AgenTRIM: Tool Risk Mitigation for Agentic AI
- garak: A Framework for Security Probing Large Language Models
- SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
- Design Patterns for Securing LLM Agents against Prompt Injections
- SynthTools: A Framework for Scaling Synthetic Tools for Agent Development
- AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications
- DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents
- A Safety and Security Framework for Real-World Agentic Systems
- Defeating Prompt Injections by Design
- MAESTRO: Multi-Agent Evaluation Suite for Testing, Reliability, and Observability
- MAPS: A Multilingual Benchmark for Agent Performance and Security
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
- Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
- DeepSeek-V3 Technical Report
- Training-Free Policy Violation Detection via Activation-Space Whitening in LLMs
- GTM: Simulating the World of Tools for AI Agents
- Identifying the Risks of LM Agents with an LM-Emulated Sandbox
- MASLab: A Unified and Comprehensive Codebase for LLM-based Multi-Agent Systems
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection