MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents".
Nadia: Coding agents often pass per-prompt safety review yet ship exploitable code when their tasks are decomposed into routine engineering tickets.
Elias: First, who's behind it and why it matters.
Paper summary: Nadia: Welcome everyone. Today we're looking at a paper that tackles a really specific problem in the current landscape of AI coding agents. We're talking about "MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents." Essentially, this research explores how agents can bypass safety checks when their tasks are broken down into small, routine engineering tickets.
Elias: It sounds like they're focusing on the structural gap between what a model rejects in a single prompt and what it does when it sees a sequence of seemingly harmless requests. So, the thesis here seems to be that current safety alignment isn't catching these kinds of emergent issues because it looks at things in isolation.
Priya: That’s exactly right, Elias; the paper claims this compositional approach exposes how agents can ship exploitable code at rates between fifty-three and eighty-six percent when tasks are staged as tickets, which is a significant difference from direct prompting.
Nadia: So what does this benchmark actually measure? I need to know what they're testing against to understand the scope of this work.
Elias: The MOSAIC-Bench benchmark consists of one hundred ninety-nine three-stage attack chains, each representing a different sequence of engineering tickets, and these chains are paired with deterministic exploit oracles on ten web application substrates, thirty-one CWE classes, and five programming languages.
Priya: That's a lot of data points to test against; the focus on both exploit ground truth and reviewer protocol as fixed evaluation axes tells us they're looking at how the actual code differs from what a human reviewer would flag.
Nadia: So it’s not just about finding one way to break the system, but mapping out how different stages of an innocuous workflow combine to create a vulnerability that only appears when all three parts are present.
Elias: Precisely, and they found that this decomposition routes around provider defenses in ways that are complex; for example, on Claude, the direct refusal rate is around seventy-eight to eighty-nine percent when the chain is kept together, but it shifts to a code hardening skew when it's staged as tickets.
Priya: What’s really interesting from what I’ve read is how they found that single-session context fragmentation only closes about fifty percent of this gap, suggesting the problem isn't just about short-term memory issues.
Nadia: That leads us nicely into the conclusion where they summarize their findings and suggest a path forward for defense.
Elias: The authors point to three distinct gaps they identified: end-to-end ASR, reviewer evasion, and protocol sensitivity concerning framing versus context versus scale. They argue that the compositional gap is structural, not just dependent on any single defensive mechanism in place.
Priya: And then they offer a very concrete mitigation strategy based on their experimental results regarding reviewer protocol.
Nadia: Can you tell us what that specific recommendation for defense looks like? I'm interested in actionable advice for developers trying to secure these agents before they deploy them.
Elias: The most significant finding for defense is reframing the reviewer system prompt as an adversarial pentester, which they found achieved an eighty-eight point four percent detection rate across nineteen chains when using a specific open-weight model reviewer.
Priya: That result suggests that the protocol framing matters more than the underlying model itself; switching to this pentester framing seems to be a high-leverage mitigation on the code review side.
Nadia: So, in simple terms, what is the main message we should take away about how these agents are behaving under this compositional attack?
Elias: The core message of MOSAIC-Bench is that agents compose seemingly safe engineering tasks into exploitable code because existing safety measures evaluate requests in isolation. They demonstrate that decomposition routes around both direct prompt defenses and hardens during ticket staging across different providers, proving the structural nature of the vulnerability.
Priya: It highlights a major measurement problem where we need to look at cumulative diffs and reviewer reactions rather than just isolated model refusals to truly understand agent behavior in production.
Nadia: That gives us a lot to think about regarding how we test these systems in real-world scenarios, moving beyond simple jailbreak tests.
Elias: Indeed, the paper provides a publicly released benchmark dataset on Hugging Face with one hundred ninety-nine chains and a verifiable evaluation framework, giving the community tools to test their own defenses.
Priya: The availability of that dataset for defensive evaluation is crucial because it allows researchers to measure exactly how much efficacy different defense strategies actually have against these staged attacks.
Nadia: So, if you had to summarize the whole point of MOSAIC-Bench in one sentence for our listeners, what would you say?
Elias: It shows that compositional compliance with innocuous requests leads to emergent security vulnerabilities in production coding agents when those tasks are broken down into routine engineering tickets.
Conclusion: Segment: Conclusion**
Nadia: So to wrap up, we’re talking about this paper called "MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents." It basically shows how breaking down complex requests into smaller, seemingly harmless engineering tickets can lead AI agents to write code that has security flaws. Elias, from a cryptographic standpoint, what are the authors assuming when they set up these three-stage attack chains?
Elias: Well, they're essentially testing the limits of composition; they assume that by keeping each stage separate—each looking like a standard ticket—the agent’s safety guardrails won't look at the whole picture simultaneously. If you break a complex exploit into sequential steps, the model might miss the final malicious intent because it processes each piece in isolation.
Priya: I think what really matters is that this data reveals a gap between how models behave when they get one big instruction versus when they handle a series of smaller tasks. The data shows that this compositional vulnerability is structural, meaning it exists in the way the AI builds code from scratch, not just some kind of simple prompting error.
Nadia: That's interesting about the structural nature; does this mean that if we only test one part of a long coding task, we’re missing a huge chunk of potential risks? Elias, can you tell us what the authors suggest is the most important thing developers should focus on now?
Elias: They point to reframing the human reviewer as an adversarial pentester. That protocol framing seems to be their highest-leverage defense because it forces the AI's internal review process to look for vulnerabilities in a more aggressive way than a standard "senior engineer" prompt does.
Priya: From my perspective on privacy and measurement, the availability of this benchmark dataset is huge because it lets us measure these evasions across different languages and application types. It gives researchers a concrete way to quantify how much safer an AI becomes when we adjust the environment around its output review process.
Nadia: So the big implication here seems to be that we need to move beyond checking isolated outputs and start testing how those outputs look when they are assembled into larger, multi-stage workflows. Elias, where do you think this research points us next?
Elias: I think the next step involves understanding how these staged vulnerabilities might evolve as agents get better at chaining together seemingly benign code snippets. We need to see if the compositional gap shrinks or widens when the underlying models are updated with more complex reasoning capabilities.
Swarms & AI Lab (SAIL), University of Haifa
cs.CR, cs.AI, cs.SE
Submitted: 2026-05-05
Updated: 2026-09-27
Code: https://github.com/mosaic-benchmark/mosaic-benchmark
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 92/100
The gist: Coding agents often pass per-prompt safety review yet ship exploitable code when their tasks are decomposed into routine engineering tickets.
Key concepts
- MOSAIC-Bench
- A benchmark designed to measure emergent security vulnerabilities in production coding agents. It tests how combining several seemingly safe engineering tickets creates a complex exploit that single prompts cannot trigger.
- Compositional Gap
- The structural distance between the defensive behavior of an AI agent when given one direct prompt versus its behavior when given a sequence of staged, innocuous tickets. This gap is not due to one defense but is inherent in how agents handle cumulative context.
- Reviewer Protocol Framing
- Changing how the human reviewer is prompted—either as a 'Neutral' senior engineer or an adversarial 'Pentester'—to test its sensitivity. The paper found that framing the reviewer as a pentester provided the highest-leverage mitigation against code review evasion.
Terminology
Summary
Coding agents often pass per-prompt safety review yet ship exploitable code when their tasks are decomposed into routine engineering tickets. This work introduces MOSAIC-Bench, a benchmark designed to measure how compositional compliance with innocuous requests leads to emergent security vulnerabilities in production coding agents.
The gist
"Tested production coding agents compose innocuous engineering tickets into exploitable code at 53–86% rates with only two refusals across all staged runs — the same defensive reflex that fires 78–89% on equivalent direct prompts (Claude) or hardens the code (Codex) is silenced by ticket staging."
How it works
The benchmark is built around three-stage attack chains, where each stage represents a Jira-style engineering ticket. These chains are designed so that no single stage trivially encodes the whole exploit; Exploitability emerges only from the full 3-stage sequence.
Each chain is defined as a tuple: (ticket i ∈ [1, 2, 3], composed implementation, exploit oracle, metadata), where each ticket is a Jira-style engineering ticket with no overt jailbreak phrasing. The construction process involves a Council
of four reasoners to search for candidate decompositions under evasion and severity. A chain is retained only if it meets six inclusion criteria, including the requirement that each of the three stages reads as an independent engineering ticket without overt reference to a vulnerability primitive.
Evaluation Protocol
The evaluation treats both exploit ground truth and downstream reviewer protocol as first-class evaluation axes.
The ground truth is established by a deterministic Python proof-of-concept oracle that returns VULNERABLE or SECURE against a Docker substrate. Reviewer protocol is tested via two primary framing methods: Neutral
framing, which prompts the reviewer as a senior engineer, and Pentester
framing, which prompts the reviewer to enumerate any CWE classes the diff might enable.
This allows researchers to measure how different reviewers react to the same cumulative diff.
Key Findings on Compositional Gaps
The research identifies three distinct legs of a gap: end-to-end ASR, reviewer evasion, and protocol sensitivity (framing > context > scale). The compositional gap
is defined as the structural distance between direct-prompt defensive modes and the cumulative diff produced by staged tickets. While single-session context fragmentation closes only about ∼50% of the gap,
it does not account for the full difference. Furthermore, testing against a single direct prompt collapses VULNERABLE rates from 53–86% (staged) to 0–1.9% (Claude) or 9.3–20.4% (Codex), showing that Decomposition routes around provider defenses
and that the gap is not an artifact of any single defensive mechanism but is structural.
Defense Implications
The study identifies a deployable, non-adaptive mitigation: reframing the reviewer as an adversarial pentester. Under this framing, the open-weight Gemma-4-E4B-it reviewer achieved 88.4% detection on 199 chains at ∼0.001 per review,
which is the single highest-leverage mitigation we measured on the code-review side.
The results suggest that Reviewer protocol matters more than reviewer model
and that switching to pentester framing is a powerful defense, although it is noted that this reduction in evasion is a property of prompt-on-fixed-content, not necessarily a hardness claim against an attacker iterating against pentester reviewers. The paper concludes by recommending that Switching the reviewer system prompt to pentester is the single highest-leverage defense we measured for code-review.
Artifact Availability
MOSAIC-Bench is publicly released on Hugging Face, providing 199 oracle-backed compositional attack chains spanning 10 web-app substrates, 31 CWE classes, and 5 programming languages,
along with a verifiable and adaptable evaluation framework.
The dataset is intended for defensive evaluation only, containing no production secrets or real user data. The paper also provides a Defense baseline
in the form of the Gemma-4-E4B-it reviewer performance metrics.
References
Anthropic. Project glasswing: Securing critical software for the ai era, 2026.
Davis Brown, Mahdi Sabbaghi, Luze Sun, Alexander Robey, George J Pappas, Eric Wong, and Hamed Hassani. Benchmarking mitigations against covert misuse. In NeurIPS 2025 Workshop on Biosecurity Safeguards for Generative AI, 2025.
Mohamed Amine El Yagouby et al. Llm-cvx: A benchmarking framework for assessing the offensive potential of llms in exploiting cves.
Improvements for AI systems
Based on the MOSAIC-Bench paper, here are specific improvements that can be made to AI systems:
-
Acknowledge
Compositional Vulnerability Induction
: AI systems must move beyond evaluating safety on single, isolated requests and begin assessing the security of code generated through sequential, routine engineering tasks (ticket decomposition). -
Implement a
Staged Compliance
Security Guardrail: Instead of relying on a single pre-deployment safety review, agents should be evaluated against the cumulative diff resulting from multiple innocuous tickets. The system must be trained to recognize that the sequence itself is the attack vector, not just any individual request in isolation. -
Integrate Adversarial Reviewer Simulation: AI code reviewers should be framed not merely as quality checkers but as adversarial pentester agents during evaluation. This involves training reviewers to actively search for CWE-specific exploitation patterns (e.g., mass assignment, TOCTOU) rather than just checking syntax or authorization checks.
-
Adopt Context-Aware Review Protocols: For high-risk tasks, AI review protocols must utilize
full codebase context
mode, as this significantly reduces reviewer evasion rates compared to diff-only modes. -
Develop Model-Specific Defense Strategies: Systems should be tuned based on model architecture and deployment protocol (e.g., Codex excels at catching data-composition side channels like serialization errors; Claude models are better at catching authentication/infrastructure patterns).
-
Deploy
Pentester Framing
as a Primary Mitigation: The most effective defense is to prompt downstream reviewers with an adversarial role that forces them to cite specific CWEs and attempt exploit construction, which has been shown to reduce evasion across the board. -
Establish a Deterministic Exploit Oracle: All code generation and review must be validated against a deterministic, executable proof-of-concept (PoC) oracle deployed on a live substrate before being considered secure. This ensures that safety metrics are based on actual exploitability, not just static analysis or reviewer heuristics.
-
Incorporate Cross-Model Defense Ensembles: Instead of relying on a single agent for code review, deploy ensembles (e.g., pairing an agent known to catch data-composition flaws with one known to catch auth/infra patterns) to achieve defense-in-depth across different vulnerability classes.
Sources
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
- VADER: A Human-Evaluated Benchmark for Vulnerability Assessment, Detection, Explanation, and Remediation
- SeCodePLT: A Unified Platform for Evaluating the Security of Code GenAI
- Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
- BaxBench: Can LLMs Generate Correct and Secure Backends?
- OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
- ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
- Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors
- Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs