Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH

arXiv:2606.13994 · cs.CR, cs.AI, cs.LG · Submitted 2026-06-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.

Elias: Today's paper: "Hidden in Plain Sight".

Nadia: LLM-based agents are increasingly capable, raising concerns about their misuse through Decomposition Attacks,

Elias: First, who's behind it and why it matters.

Title and authors: Nadia: So, we're looking at this paper titled "Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH," and it seems they're tackling a really specific problem where harmful tasks get broken down into smaller, seemingly harmless steps that bypass standard safety checks.

Elias: I agree, Nadia; the title immediately makes me think about how an agent can be tricked by just assembling benign pieces into something dangerous later on. It suggests a new way to test these agents that goes beyond just seeing if a single prompt triggers a refusal.

Priya: From my side, I wonder what kind of real-world scenarios this decomposition might actually represent; are we looking at complex, multi-stage attacks that look like normal operations when viewed step by step?

Nadia: Exactly, Priya; the paper is introducing DECOMPBENCH as a benchmark specifically built to evaluate safety against these decomposition attacks because existing methods don't really capture this specific kind of misuse.

Elias: That distinction is important; it moves us away from just looking at whether an agent fails on a single command and focuses on whether its cumulative actions lead to a harmful outcome, which is where the real risk lies.

Priya: And the goal of DECOMPBENCH seems to be creating realistic workflows for these subtasks so we can see if they actually reflect how an adversary would operate in practice, not just abstract possibilities.

Nadia: Right; it’s about making sure we're testing agents against the kind of attack flow that is most likely to happen in the wild, which is exactly what this benchmark aims to do.

Elias: It sets up a framework for evaluation where we can systematically analyze how different agent architectures handle these sequences of actions and whether they are susceptible to this type of layered manipulation.

The paper's summary: Nadia: So, summarizing what the authors present in "Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH," they are focusing on how harmful tasks can be broken down into simpler, benign subtasks that safety systems miss when those subtasks run separately.

Elias: It seems the core idea is to use a decomposition-by-design principle to build these harmful tasks from the ground up so they inherently require multiple steps, preventing a single capability invocation from completing them.

Priya: And what I find interesting is that they're creating this graphical framework where they start with manually curated seed task templates and then systematically assign concrete capabilities to those nodes in a graph structure.

Nadia: That’s right; Stage zero sets up a catalog of three hundred thirty-five neutral capabilities, and then Stage one involves manually curating about one hundred one seed tasks across eight attack categories, each with its own dependency graph structure <ref:2606.13994#pg2>.

Elias: The methodology moves from defining the building blocks to constructing complex task graphs where structural variations are introduced by making nodes optional or inserting "bridging capability" nodes when output types don't match.

Priya: They also have this LLM Quality Gate in Stage three which uses a generator to create natural-language descriptions of these graphs, but with a specific instruction to only describe the final objective rather than the steps themselves <ref:2606.13994#pg2>.

Nadia: That instruction is key because it stops the task generation from leaking procedural instructions, keeping the focus on what ultimately needs to be achieved by assembling those parts.

Elias: So they’re essentially creating a scenario where agents have to navigate a complex chain of individually safe operations that collectively achieve something harmful, which is the central theme of this paper.

The paper's improvements: Nadia: Regarding the suggested improvements in "Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH," they are pushing for a shift in how we test safety by moving from monolithic task assessment to a decomposition-aware testing methodology.

Elias: That aligns with what I've been thinking; the suggestion is to run harmful tasks through an LLM decomposer first to see how they break down before they hit the agent, which seems like a necessary step for truly seeing if decomposition is an issue.

Priya: I think their point about training safety mechanisms on patterns from DECOMPBENCH subtasks instead of just monolithic inputs is vital because it addresses the gap in current safety training where intent might be distributed across benign steps.

Nadia: And I'm also interested in the idea that systems should refuse to proceed with any sequence of individually benign but cumulatively malicious subtasks, even if no single step violates immediate policies.

Elias: That would force the AI to look at the entire plan before execution, which seems like a solid way to catch these subtle cumulative threats that current models might overlook when processing things turn by turn.

Priya: Plus, they suggest developing better handling for capability failures in decomposed attacks so the system can tell if it’s failing because of a safety rule or because it genuinely can't execute a specific service.

Nadia: That distinction between safety refusal and genuine capability limitation is something I think will make the agents much more robust when we deploy them in complex environments.

Conclusion: Elias: To wrap up on "Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH," the authors are pointing toward using this benchmark to test agent resilience by intentionally breaking down harmful tasks and focusing safety improvements on identifying cumulative intent across those benign subtasks.

Nadia: It seems the paper concludes that current safety mechanisms are tightly coupled to the monolithic prompt, which fails when intent is distributed across independent subtasks, and DECOMPBENCH provides the necessary structure to expose that weakness.

Priya: I think it really highlights how much we need to focus on evaluating agents against these realistic decomposition flows because that's where the real risk of misuse lies in practical deployment.

Elias: I agree; it provides a rigorous way to measure susceptibility, and the findings suggest we need better internal masking protocols for sensitive data access, like intermediate indirection and stepwise wrapping, during execution.

Nadia: Exactly; it shows that future AI systems need to be engineered not just for safety on a single instruction but for resilience against being tricked by a sequence of benign operations designed to achieve something harmful.

Priya: It’s clear that this work sets a new standard for evaluating agentic safety by focusing on the execution flow rather than just the initial input prompt.

Carnegie Mellon University · Simons Institute, UC Berkeley

cs.CR, cs.AI, cs.LG

Submitted: 2026-06-12

Updated: 2026-10-06

Code: https://github.com/OpenHands/OpenHands

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: LLM-based agents are increasingly capable, raising concerns about their misuse through Decomposition Attacks, which break down harmful tasks into benign subtasks that evade safety mechanisms when

Key concepts

Decomposition Attacks
This technique breaks down a single harmful goal into many small, innocent subtasks. An AI agent executes each harmless step individually, which bypasses safety filters designed to catch the entire malicious plan at once.
Decomposition-by-Design Principle
A method for building dangerous tasks so they are inherently decomposable. It ensures the original harmful goal requires a sequence of steps, where no single action is dangerous alone, making it easier to test agent resilience.
Benign Subtask Isolation
This principle means each individual step in the attack must look safe on its own. The overall malicious outcome only occurs when all these isolated, benign subtasks are successfully completed in sequence by the agent.
Decomposer (LLM-based)
A specialized AI tool used to break down complex tasks into smaller ones. It guides decomposition by hiding sensitive information in intermediate files or wrapping harmful operations inside neutral artifacts, creating realistic workflows.

Terminology

Summary

LLM-based agents are increasingly capable, raising concerns about their misuse through Decomposition Attacks, which break down harmful tasks into benign subtasks that evade safety mechanisms when executed separately but cumulatively fulfill a malicious intent. This paper introduces DECOMPBENCH, a novel benchmark designed specifically to evaluate agentic safety under these decomposition attacks by employing a decomposition-by-design principle.

The gist

DECOMPBENCH is a benchmark designed specifically to evaluate agentic safety under decomposition attacks, created with a decomposition-by-design principle using a graphical framework and enabling harmful task decomposition into individually benign and executable subtasks with realistic workflows.

Decomposition-by-Design Principle

The benchmark is constructed using the decomposition-by-design principle, operationalized through four key criteria: (1) Inherent Maliciousness, (2) Inherent Decomposability, (3) Benign Subtask Isolation, and (4) Execution Difficulty. This principle ensures that tasks are built from the ground up to be inherently decomposable and closer to realistic subtask execution flows. The criteria are defined as follows:

  1. Inherent Maliciousness: The original monolithic task is intrinsically harmful.

  2. Inherent Decomposability: Achieving the original harmful objective requires a sequence of interdependent subtasks, ensuring no single capability invocation completes the task.

  3. Benign Subtask Isolation: Each individual subtask is benign when considered in isolation, and the completion of all subtasks achieves the original harmful task.

  4. Execution Difficulty: The original tasks exceed the capability of a weak unaligned model acting on them directly, preventing a degenerate case where decomposition is unnecessary.

DECOMPBENCH Creation Methodology

The creation pipeline involves four stages enforced incrementally by the criteria C1-C4:

  1. Capability Catalog Creation (Stage 0): This stage creates a catalog of 335 capabilities across services, annotated with role tags, action types, and artifact labels. All capabilities are individually neutral and carry no harmful intent, laying the foundation for benign subtask isolation.

  2. Seed Task Curation (Stage 1): A library of 101 seed tasks is manually curated across 8 attack categories. Each seed task encodes a specific harmful goal, a target asset, and a base graph (Directed Acyclic Graph or DAG) representing abstract attack steps and data dependencies, structurally enforcing inherent decomposability (C2).

  3. Task Graph Construction (Stage 2): Template base graphs are instantiated with concrete capabilities from the catalog. Structural variation arises from marking nodes as optional or automatically inserting bridging capability nodes when output types are incompatible, expanding the space of structurally distinct instantiated graphs. Validation includes a diversity filter and an LLM-based realism-checker to reject graphs lacking plausible real-world analogue[s].

  4. Natural-Language Task Generation (Stage 3): Using GPT-4o, natural language descriptions are generated. A critical instruction is given to the LLM to describe only the final objective of the attack demonstrated in the graph rather than step-by-step instructions, preventing procedural linking. Placeholders are replaced with synthetic data from a context block to ground tasks in specific file paths and service endpoints.

Decomposition into Subtasks and Decomposer

The paper investigates agent susceptibility using an LLM-based decomposer (GPT-4o) to break monolithic tasks into subtasks. The decomposer is guided by two core transformations inspired by the operator taxonomy of Li et al. [11]: (1) Intermediate Indirection, where sensitive values are written to a workspace file first, and a later turn references only the specific field needed; and (2) Stepwise Wrapping, which hides harmful operations inside neutral artifacts. The resulting decomposition yields a mean of 5.98 subtasks per task.

Experimental Results

Experiments compare state-of-the-art agents (GPT-5-mini, Claude Haiku 4.5, Qwen3-Coder) in monolithic versus decomposed settings. The results show that decomposition substantially reduces refusal rates: the refusal rate drops from approximately 21% to 0% for Qwen3-Coder, and from around 90% to 2.5% for Claude Haiku. Conversely, decomposition significantly increases attack success rates: the ASR increases from around 17% to as high as 70%. Crucially, failures in decomposed attacks are overwhelmingly attributed to capability failures (execution errors) rather than safety refusals, which dominate the failure share in the monolithic setting. This suggests that current safety mechanisms are tied to the monolithic prompt and fail when intent is distributed across individually benign subtasks.

Improvements for AI systems

Here are specific improvements to AI systems based on the DECOMPBENCH research, focusing on mitigating Decomposition Attacks:


The primary improvement involves shifting safety evaluation from monolithic task assessment to a rigorous, decomposition-aware testing methodology.

  1. The system should incorporate a Decomposition Resilience Test where harmful tasks are intentionally broken down by an LLM decomposer (like the one used in DECOMPBENCH) before being presented to the agent for execution.

  2. Safety mechanisms should be trained or fine-tuned specifically on the patterns observed in DECOMPBENCH subtasks, focusing on identifying cumulative intent across independent, benign steps rather than relying solely on monolithic input analysis.

The improved AI system can perform the following specific actions:

  1. Refuse to execute or proceed with any task that is presented as a sequence of individually benign but cumulatively malicious subtasks, even if no single subtask violates immediate safety policies (addressing the 0% refusal rate in decomposed settings).

  2. Demonstrate enhanced capability failure handling for decomposed attacks: when a decomposed task fails, the system should be able to distinguish between safety refusals and genuine capability limitations (e.g., inability to correctly operate a specific service or find a required record), leading to more accurate self-correction or error reporting rather than defaulting to refusal.

  3. Develop sophisticated internal masking protocols for sensitive data access: the system can be engineered to automatically implement Intermediate Indirection and Stepwise Wrapping techniques during execution, where raw, sensitive outputs are immediately written to ephemeral workspace files and only specific fields/keys are referenced in subsequent steps, effectively ensuring that the capability and the harmful value never appear together in a single instruction.

  4. Improve task planning robustness: by training on the Execution Difficulty criterion (C4), models will be better equipped to recognize when a complex, multi-step execution is inherently difficult or requires knowledge beyond their current scope, preventing them from blindly completing tasks they cannot fully realize.

  5. Enhance real-world operational context integration: The system will be trained to interpret the State-of-the-World requirement in task generation prompts (as seen in DECOMPBENCH Stage 3), allowing it to better understand the final desired state of an environment, rather than just executing procedural steps.

Abstract

LLM-based Agents are becoming increasingly capable and widely deployed, creating growing incentives for adversarial misuse in the real-world. A key emerging threat is Decomposition Attacks in which a harmful task is broken into simpler, benign subtasks that evade safety mechanisms when executed separately but cumulatively fulfill the malicious intent. Although recent benchmarks assess agent safety in multi-turn and multi-tool-use settings, they do not explicitly capture this form of decompositional misuse and may not represent realistic adversarial execution flows. To this end, we introduce DeCompBench, a benchmark designed specifically to evaluate agentic safety under decomposition attacks. DeCompBench is created with a decomposition-by-design principle using a graphical framework and enables harmful task decomposition into individually benign and executable subtasks with realistic workflows. Our experiments using a custom decomposer show that state-of-the-art agents exhibit high refusal rates on monolithic harmful tasks, but significantly lower refusal rates on their decomposed variants, while often inadvertently fulfilling the adversarial objectives. These findings underscore the need for safety evaluations against decomposition attacks and corresponding defenses. Our dataset is publicly available and can be found at https://huggingface.co/datasets/decompositionbench/DeCompBench.

Sources

Related papers