CART: Closed-Loop Adaptive Red Teaming for Large Language Models
cs.AI
Submitted: 2026-09-23
Updated: 2026-09-23
Terminology
Sources
- Technical Report: Evaluating Goal Drift in Language Model Agents
- Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO
- Jailbreaking Black Box Large Language Models in Twenty Queries
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Taming Various Privilege Escalation in LLM-Based Agent Systems: A Mandatory Access Control Framework
- SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
- Lessons From Red Teaming 100 Generative AI Products
- Red Teaming Language Models with Language Models
- Ignore Previous Prompt: Attack Techniques For Language Models
- Identifying the Risks of LM Agents with an LM-Emulated Sandbox
- MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers
- Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
- GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection