Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
cs.CR, cs.AI
Submitted: 2025-10-01
Updated: 2026-08-30
Comments: 22 pages, 18 figures, 8 tables
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Program Synthesis with Large Language Models
- CodeT: Code Generation with Generated Tests
- Evaluating Large Language Models Trained on Code
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs