Toward Secure Code Generation: Bridging Correctness and Security via Task-Adaptive Vulnerability Modeling and Execution-Based Benchmarking
cs.SE, cs.CL
Submitted: 2024-07-02
Updated: 2026-09-14
Comments: Substantially extended and revised version. CodeSecEval has been expanded from 180 tasks covering 44 vulnerability types to 255 executable Python tasks spanning 77 CWE categories. This version additionally introduces SecAwareCoder, a task-adaptive threat-modeling and execution-guided framework for secure code generation, together with new experiments and analyses. Accepted/published in ISSTA 2026
Code: https://github.com/github/codeq
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) are increasingly used for program synthesis, yet they often generate code that is functionally plausible but insecure.
Terminology
Abstract
Large language models (LLMs) are increasingly used for program synthesis, yet they often generate code that is functionally plausible but insecure. Progress in secure code generation has been hindered by benchmarks that are small, non-executable, leak mitigation details, or rely on noisy analyzers and subjective judgments, making it difficult to measure whether security improves without sacrificing correctness. We address these gaps with CodeSecEval, an execution-based benchmark for secure code generation, comprising 255 Python tasks spanning 77 CWE categories. Each task provides paired insecure and secure implementations together with executable functional and vulnerability-targeted security tests, enabling precise and reproducible evaluation of secure code generation and insecure-code repair. Building on CodeSecEval, we propose SecAwareCoder, an agent-based framework that shifts code generation toward secure-by-construction synthesis. SecAwareCoder performs task-adaptive threat modeling to identify security-sensitive regions and derive task-grounded vulnerability hypotheses, uses these hypotheses to guide both constraint-aware code generation and security-aware test synthesis, and leverages execution feedback for targeted refinement. Experiments across multiple LLM backbones show that SecAwareCoder consistently improves Pass@1 and security robustness over prompting and analyzer-driven baselines, narrowing the security--correctness gap in LLM code generation.
Sources
- Program Synthesis with Large Language Models
- Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models
- Evaluating Large Language Models Trained on Code
- PaLM: Scaling Language Modeling with Pathways
- CodeBERT: A Pre-Trained Model for Programming and Natural Languages
- InCoder: A Generative Model for Code Infilling and Synthesis
- Self-planning Code Generation with Large Language Models
- How Secure is Code Generated by ChatGPT?
- DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation
- StarCoder: may the source be with you!
- Think Outside the Code: Brainstorming Boosts Large Language Models in Code Generation
- ReACC: A Retrieval-Augmented Code Completion Framework
- Comparing Software Developers with ChatGPT: An Empirical Investigation
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis
- GPT-4 Technical Report
- Do Users Write More Insecure Code with AI Assistants?
- Code Llama: Open Foundation Models for Code
- Automatic Static Bug Detection for Machine Learning Libraries: Are We There Yet?
- LLaMA: Open and Efficient Foundation Language Models
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties