ALIBI: Adaptive Agentic Attacks on LLM-Based Vulnerability Detectors via Adversarial Code Comments
Zixuan Wu, Cristina Nita-Rotaru
cs.CR
Submitted: 2026-07-27
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models are increasingly deployed for security-sensitive tasks such as vulnerability detection and code review.
Terminology
Abstract
Large language models are increasingly deployed for security-sensitive tasks such as vulnerability detection and code review. Their reliance on natural-language context embedded in source code exposes a previously underexplored attack surface: adversarial comments that can influence a detector's reasoning without changing program behavior. We study LLM-based vulnerability detectors against a new adversary: a coding agent that implements new functionality, deliberately introduces vulnerabilities, and strategically inserts adversarial source-code comments to evade detection. We present ALIBI, an automated adaptive black-box attack framework that generates and iteratively refines adversarial comments using detector reasoning and feedback. We transform real-world vulnerability-fixing commits into coding tasks and evaluate four representative LLM-based vulnerability detectors, ranging from specialized open-weight reasoning models to frontier multi-agent systems. All evaluated detectors are highly vulnerable: attack success rates exceed 90% across 125 real-world null-pointer dereference vulnerabilities, reaching 100% on one system. The framework also generalizes beyond this vulnerability class. Adversarial comments steering detector reasoning or fabricating external tool results prove most effective, while iterative refinement based on detector feedback further increases attack success. Finally, prompt-level defenses provide limited robustness against adaptive attacks, whereas architectural isolation and pre-detector comment sanitization substantially improve resilience. Our findings expose a fundamental attack surface in current LLM-based vulnerability detectors and motivate security-aware designs that carefully calibrate trust between natural-language context and program evidence.
Sources
- VulWeaver: Weaving Broken Semantics for Grounded Vulnerability Detection
- Securing AI Agents with Information-Flow Control
- Defeating Prompt Injections by Design
- Reducing False Positives in Static Bug Detection with LLMs: An Empirical Study in Industry
- Minimizing False Positives in Static Bug Detection via LLM-Enhanced Path Feasibility Analysis
- LLM-based Vulnerability Detection at Project Scale: An Empirical Study
- PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
- CoTDeceptor:Adversarial Code Obfuscation Against CoT-Enhanced LLM Code Agents
- From SFT to RL: Demystifying the Post-Training Pipeline for LLM-based Vulnerability Detection
- Everything You Wanted to Know About LLM-based Vulnerability Detection But Were Afraid to Ask
- RepoAudit: An Autonomous LLM-Agent for Repository-Level Code Auditing
- IRIS: LLM-Assisted Static Analysis for Detecting Security Vulnerabilities
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- LLM Can be a Dangerous Persuader: Empirical Study of Persuasion Safety in Large Language Models
- "Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills in the Wild
- When Names Disappear: Revealing What LLMs Actually Understand About Code
- Backdoors in Neural Models of Source Code
- Agentic Much? Adoption of Coding Agents on GitHub
- You Autocomplete Me: Poisoning Vulnerabilities in Neural Code Completion
- Measuring and Exploiting Contextual Bias in LLM-Assisted Security Code Review
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs