Vulnerable Code Search: Transferable Attack for Code Language Models
cs.SE, cs.CR
Submitted: 2026-08-26
Updated: 2026-08-26
Code: https://github.com/salesforce/CodeT5
Project page: https://archersama.github.io/coir
Terminology
Sources
- Evaluating Large Language Models Trained on Code
- CodeBERT: A Pre-Trained Model for Programming and Natural Languages
- OASIS: Order-Augmented Strategy for Improved Code Search
- Qwen2.5-Coder Technical Report
- Explaining and Harnessing Adversarial Examples
- GraphCodeBERT: Pre-training Code Representations with Data Flow
- Large Language Model (LLM) for Software Security: Code Analysis, Malware Analysis, Reverse Engineering
- Adversarial Attacks on Code Models with Discriminative Graph Patterns
- StarCoder: may the source be with you!
- RepoQA: Evaluating Long Context Code Understanding
- StarCoder 2 and The Stack v2: The Next Generation
- Code Llama: Open Foundation Models for Code
- SynCoBERT: Syntax-Guided Multi-Modal Contrastive Pre-Training for Code Representation
- Qwen2 Technical Report
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation
- Learning Deep Semantic Model for Code Search using CodeSearchNet Corpus
- Repoformer: Selective Retrieval for Repository-Level Code Completion
- Retrieval-Augmented Generation for AI-Generated Content: A Survey
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties