Specification Before Generation: A Pre-Registered, Five-Model Paired Evaluation of a Specification Frame for LLM-Generated Code in Money, Time, Idempotency, and Access Tasks
cs.SE, cs.CR
Submitted: 2026-09-20
Updated: 2026-09-30
Comments: 10 pages plus references, 2 figures, 5 tables. Dataset v1.1: https://doi.org/10.5281/zenodo.22850887. Also deposited at https://doi.org/10.5281/zenodo.22830263
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- How secure is AI-generated Code: A Large-Scale Comparison of Large Language Models
- Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
- ContextBench: A Benchmark for Context Retrieval in Coding Agents
- The Productivity-Reliability Paradox: Specification-Driven Governance for AI-Augmented Software Development
- Understanding Specification-Driven Code Generation with LLMs: An Empirical Study Design
- Specifications: The missing link to making the development of LLM systems an engineering discipline
- Towards Specification-Driven LLM-Based Generation of Embedded Automotive Software
- Broken by Default: A Formal Verification Study of Security Vulnerabilities in AI-Generated Code
- Constitutional Spec-Driven Development: Enforcing Security by Construction in AI-Assisted Code Generation
- Is Your Prompt Poisoning Code? Defect Induction Rates and Security Mitigation Strategies
- On Fixing Insecure AI-Generated Code through Model Fine-Tuning and Prompting Strategies
- RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties