Proteus: A Self-Evolving Red Team for Agent Skill Ecosystems
cs.CR, cs.AI
Submitted: 2026-05-12
Updated: 2026-09-26
Code: https://github.com/microsoft/PyRIT
Terminology
Sources
- Constitutional AI: Harmlessness from AI Feedback
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents
- SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- "Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills in the Wild
- Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills
- GAIA: a benchmark for General AI Assistants
- AgenticRed: Evolving Agentic Systems for Red-Teaming
- Active Attacks: Red-teaming LLMs via Adaptive Environments
- Genesis: Evolving Attack Strategies for LLM Web Agent Red-Teaming
- AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs