Detecting Data Poisoning in Code Generation LLMs via Black-Box, Vulnerability-Oriented Scanning
cs.CR, cs.AI, cs.SE
Submitted: 2026-03-17
Updated: 2026-09-24
Terminology
Sources
- GPT-4 Technical Report
- Gradient-based Adversarial Attacks against Text Transformers
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- Qwen2.5-Coder Technical Report
- The Stack: 3 TB of permissively licensed source code
- A Systematic Study of Code Obfuscation Against LLM-based Vulnerability Detection
- Code Llama: Open Foundation Models for Code
- StarCoder 2 and The Stack v2: The Next Generation
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Universal Adversarial Triggers for Attacking and Analyzing NLP
- BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs