Agent Hacks Agents: Autoresearch Discovers Vulnerabilities in Production Agents
cs.CR, cs.AI
Submitted: 2026-07-13
Updated: 2026-09-26
Code: https://github.com/henrymao2004/Auto-research-red-teaming-in-sleep
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- AgentRx: Diagnosing AI Agent Failures from Execution Trajectories
- CausalFlow: Causal Attribution and Counterfactual Repair for LLM Agent Failures
- DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents
- IterInject: Indirect Prompt Injection Against LLM Agents via Feedback-Guided Iterative Optimization
- RainbowPlus: Enhancing Adversarial Prompt Generation via Evolutionary Quality-Diversity Search
- AJAR: Adaptive Jailbreak Architecture for Red-teaming
- How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
- AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
- AutoRISE: Agent-Driven Strategy Evolution for Red-Teaming Large Language Models
- A Comprehensive Survey on Benchmarks and Solutions in Software Engineering of LLM-Empowered Agentic System
- Automated Hypothesis Validation with Agentic Sequential Falsifications
- T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
- AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
- ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
- ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs
- CHASE: Adversarial Red-Blue Teaming for Improving LLM Safety using Reinforcement Learning
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- David vs. Goliath: Verifiable Agent-to-Agent Jailbreaking via Reinforcement Learning
- Measuring Agents in Production
- Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs