Adversarial Prompting Framework for AI Safety Assessment
Yash Bhatnagar, Kunal Banerjee, Anirban Chatterjee
cs.CR, cs.AI
Submitted: 2026-07-15
Comments: 3 pages, 1 figure, presented as a poster at International Conference on Data Science (CODS), December 17-20, 2025, Pune, India
Code: https://github.com/unitaryai/detoxify
License: http://creativecommons.org/licenses/by/4.0/
The gist: Artificial Intelligence (AI), especially Generative AI (GenAI), adoption has increased in industries significantly in recent years.
Terminology
Abstract
Artificial Intelligence (AI), especially Generative AI (GenAI), adoption has increased in industries significantly in recent years. However, the use of these models may also expose systems to new forms of cyberattacks by different malicious actors -- adversarial prompt attack (APA) being one of the most prominent examples of such threats. This paper presents the implementation of an Adversarial Prompting Framework (APF) for a comprehensive assessment of AI safety. The framework systematically evaluates the resilience of the AI model through the generation of structured adversarial prompts at multiple sophistication levels, from direct harmful requests to advanced encoding-based attacks. Our implementation demonstrates the practical application of this methodology in enterprise environments, providing automated testing capabilities with quantitative security assessment metrics. The results indicate significant variations in the model vulnerabilities across different attack vectors, with encoded prompts presenting the highest success rates in bypassing safety mechanisms.
Sources
- Securing Large Language Models: Threats, Vulnerabilities and Responsible Practices
- Jailbreak Transferability Emerges from Shared Representations
- Design Patterns for Securing LLM Agents against Prompt Injections
- Recent Advances in Attack and Defense Approaches of Large Language Models
- SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
- Security Concerns for Large Language Models: A Survey
- Large Language Models for Cyber Security: A Systematic Literature Review
- Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs