Defensive Sufficiency in a Stackelberg Model of AI Security
cs.CR, cs.AI
Submitted: 2026-10-07
Updated: 2026-10-08
Code: https://github.com/rajlakshmichavan/garak-in-lean
Terminology
Sources
- garak: A Framework for Security Probing Large Language Models
- Toward a Dynamic Stackelberg Game-Theoretic Framework for Agentic AI Defense Against LLM Jailbreaking
- Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
- PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System
- Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs