Trajectory-Level Security Debt in LLM Coding Agents
cs.CR, cs.SE
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/ProgramBench/submissions
Terminology
Sources
- Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models
- Why Do Multi-Agent LLM Systems Fail?
- Evaluating Large Language Models Trained on Code
- Teaching Large Language Models to Self-Debug
- NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
- Shift-Up: A Framework for Software Engineering Guardrails in AI-native Software Development -- Initial Findings
- Proximal Policy Optimization Algorithms
- BaxBench: Can LLMs Generate Correct and Secure Backends?
- DeceptPrompt: Exploiting LLM-driven Code Generation via Adversarial Natural Language Instructions
- ProgramBench: Can Language Models Rebuild Programs From Scratch?
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs