Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems
cs.CR, cs.AI
Submitted: 2026-06-18
Updated: 2026-09-30
Code: https://github.com/llm-attacks/llm-attacks
Terminology
Sources
- Deliberative Alignment: Reasoning Enables Safer Language Models
- Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
- "To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios
- Hacking Back the AI-Hacker: Prompt Injection as a Defense Against LLM-driven Cyberattacks
- GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs