Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
summary
The gist
The gist The study measured what changes when security requirements, selected from a governed security-by-design knowledge base (SbD-ToE) and delivered through a Model Context Protocol (MCP) server,
In short
The study tested whether giving AI code generators specific security requirements at execution time improves code security. By delivering requirements from a governed knowledge base via a Model Context Protocol, researchers found that this delivery significantly raised security scores on two benchmarks. The finding suggests that the gap in secure-code benchmarks often stems from missing specifications rather than inherent model weakness.
Key concepts
- Security-by-Design Knowledge Base (SbD-ToE)
- This is a governed knowledge base containing security requirements for applications. These requirements are structured and versioned, linking them to standards like OWASP ASVS and NIST SSDF. It acts as the source from which specific security rules are drawn.
- Model Context Protocol (MCP) Server
- The MCP server is a mechanism used to expose the SbD-ToE knowledge base to coding agents. It uses a selector tool to match task details with requirement families, returning the exact set of activated controls and objectives needed for a specific task.
- Execution Fidelity (MSbD)
- This proposed measure aims to separate delivery failure from implementation failure. It would count how many applicable security requirements the generated code successfully satisfies, providing a more accurate assessment than simple pass/fail metrics.
Terminology used across episodes
This episode discusses
- Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code · Paper Radio
- DualGauge: Automated Joint Security-Functionality Benchmarking of Specification-Only Code Generation by LLMs and Coding Agents
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions
- CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation
- SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
- BaxBench: Can LLMs Generate Correct and Secure Backends?
- Instruction Tuning for Secure Code Generation
- RESCUE: Retrieval Augmented Secure Code Generation
- Activating Latent Security Knowledge through LLM-Guided Risk Analysis for Secure Code Generation · Paper Radio
- AutoSafeCoder: A Multi-Agent Framework for Securing LLM Code Generation through Static Analysis and Fuzz Testing
- Multi-Agent Code-Orchestrated Generation for Reliable Infrastructure-as-Code
- Large Language Models are not Fair Evaluators
The paper
Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code · Read on arXiv
Pedro Farinha
Security by design asks that security requirements are defined before code is written. Secure-code benchmarks typically measure the opposite situation: the agent receives the task without requirements and is scored with security tests it has not seen. We measured what changes when security requirements, selected from a governed security-by-design knowledge base (SbD-ToE) and delivered through a Model Context Protocol (MCP) server, are given to the generator at the point of execution. On DualGauge, which scores with a language-model judge, we ran 59 Python tasks: the share of tasks passing all security tests rose from 44.1% to 78.0% and the share of security tests passed from 77.4% to 93.0%. On BaxBench, which runs functional tests and real exploits in containers, we ran 28 backend scenarios: among functionally correct solutions, the share with no successful exploit rose from 65% to 86%, comparable to the authors' Oracle Security Reminder (85%), which they call an unrealistic upper bound; the requirement delivery reached it without knowing the tests. The study has limits. The pre-specified joint secure-and-functional metric did not improve significantly on either benchmark. The code written with the governed requirements is stricter, and it failed functional tests that fix values the task never states; we attribute those failures case by case in a post-hoc analysis. Each experiment used one model and one generation per task. The requirements come from the author's open-source manual; we do not assess its completeness. Governed security requirements delivered at the point of execution appear to improve the security of generated code. Current benchmarks cannot tell whether that code meets the security requirements that apply to it, only whether it shows the weaknesses their evaluators were built to detect. We propose measuring conformance with the applicable requirements.
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Applying Security by Design at the Point of Execution".
Elias: The gist The study measured what changes when security requirements, selected from a governed security-by-design knowledge base (SbD-ToE) and delivered through a Model Context Protocol (MCP) server,
Nadia: First, who's behind it and why it matters.
Paper summary: Nadia: So to wrap up this discussion on "Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code," what's the big picture takeaway here? It’s not that we need a completely new AI architecture to fix security.
Elias: It’s more about how we structure the input pipeline. The paper shows that delivering specific, governed security requirements through a mechanism like the Model Context Protocol can raise the security of AI-generated code even when you keep the model exactly the same.
Priya: The main implication for people who only listen to this is that when you see a gap in secure code benchmarks, stop immediately and ask if those requirements were delivered to the generator or if they were just assumed knowledge.
Nadia: That addresses that specification gap we talked about earlier. It suggests that improving security isn't always about fixing the model's brain; sometimes it’s about improving the instructions it gets when it starts working.
Elias: The authors are suggesting a process where an organization decides which security requirements apply to an application, and then that set of rules is given to the AI at execution time by a server. It brings security back into that design phase before the code is even written.
Priya: They flag a specific limitation in their work, though. They found that on both benchmarks, the pre-specified joint secure-and-functional metric didn't improve significantly at all when they compared it to the baseline without requirements.
Nadia: That means while adding requirements helps in certain scenarios, we can't just rely on a single combined score to tell us everything is better. We still need those functional tests and other checks for a full picture.
Elias: Exactly. The study shows that this requirement delivery mechanism is a tool you can use in your workflow to make the AI output more secure, but it isn't a magic fix for every security problem out there.
Priya: So the future work they propose focuses on creating a measure of conformance, something that counts how many of those specific requirements the code actually satisfies. That moves us closer to measuring specification adherence directly.
Conclusion: Nadia: So we’re wrapping up on this study about delivering security requirements to AI right when it starts generating code and what that actually means for building software.
Elias: The paper, "Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code," is looking at how giving an AI specific security rules changes the code it produces.
Priya: It’s essentially testing if you can inject a set of known security requirements into the AI's execution process without changing the underlying model itself, and whether that helps make the output safer.
Nadia: Exactly. The authors are showing that this delivery mechanism, using something like a Model Context Protocol server, does raise security metrics on benchmarks they tested.
Elias: But Priya pointed out a key caveat in their results—the joint metric didn't always improve significantly on either benchmark. That means we can't just look at one combined score to say everything is better than it was before.
Priya: That makes sense because they ran experiments with fixed models and tasks, and the functional tests sometimes failed because the code was structurally correct but didn't meet a specific requirement that wasn't in the task itself.
Nadia: So, what does this mean for us on a daily basis? It suggests that when we have known security rules for an application, making sure those rules are actively fed to the AI at runtime could be a real way to close specification gaps.
Elias: It shifts the focus from just hoping the model is good enough to actually giving it the necessary constraints right where it’s building something.
Priya: And for someone who only listens, it means that instead of just blaming the AI for writing insecure code, we might need to focus on whether those specific security rules were actually delivered or not.
Nadia: That’s what they are proposing next: a measure of how well the generated code conforms to those specific requirements.
Elias: It sounds like the next step is moving beyond just saying "the code is secure" to proving "the code satisfies these specific constraints."
Priya: And that leads into their suggestion for this new measure, execution fidelity, which tries to separate whether the AI failed to give you a requirement versus whether it just implemented the wrong logic.
Nadia: We’ll have to see if that kind of measurement actually proves useful in real-world scenarios beyond these specific benchmark tests.
More episodes
- 2610.10597-Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive Rigging
- 2610.10608-From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage
- 2610.10612-PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits
- 2610.10742-BRANCH: Bypassing Multi-Scanner AI Guardrails