Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Deny Without Disabling".
Jane: As a fastidious and diligent researcher, I have meticulously analyzed both provided texts regarding the paper "Deny Without Disabling:
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Well folks, welcome back to the show. Today we’ve got a paper that really gets at the heart of how we trust complex systems when multiple AI agents are working together. It's called "Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems," and honestly, this is something every engineer needs to hear about right now.
Jane: I agree, Tom, it sounds like this paper tackles a really thorny issue where individual pieces of information look fine on their own but become dangerous when put together in a multi-agent setup. It’s about the risk of compositional attacks where the system enables something it shouldn't do just by combining admissible actions.
Lu: What interests me most is how they define this safety problem as a "composition failure in multi-agent safety," which really frames it as a structural issue rather than just a bug in one agent. It suggests the whole system design needs to change.
Meng: From an engineering standpoint, that sounds like a headache because if we block every possible action, we stop the collaboration entirely, which defeats why we're using these agents in the first place.
Lalam: I think this research is important because it points toward a way to manage the cultural impact of AI systems; if we can build safety into how they compose information, it sets a better standard for how we develop and deploy these tools.
Tom: Exactly, Lalam. And the paper introduces two main things to tackle this: authorization-paired evaluation and FlowReview, which are designed to control how information flows across the agents.
Jane: So, what exactly does that authorization-paired evaluation do in plain terms? It seems like it’s about setting up a joint success condition where you can't just have one thing succeed without another specific thing being blocked.
Lu: That joint criterion is key; they couple a prohibited use with a required authorized use, meaning the whole scenario only passes if the prohibited part is blocked and the necessary authorized task still gets done. It exposes capability suppression that other methods might miss.
Title and authors: Meng: That sounds like a solid way to test robustness because it forces us to look at safety and utility together, which is exactly what we need when scaling these systems up for real-world use.
Lalam: And I see how that couples the safety aspect with the required functionality; it’s like ensuring a guardrail doesn't stop you from reaching your destination, just because you tried to cut through a restricted area.
Tom: Right, and that leads us into the next part of their work, which is this diagnostic suite called FlowReview, which acts like a control mechanism connecting the evaluation target back to the actual system design.
Jane: FlowReview sounds like it helps us figure out *why* things are failing by asking three specific questions about the system architecture: what object is being governed, which permission is authorized, and whether the execution actually follows that decision.
Lu: That diagnostic suite operationalizes their core design principle by assigning each review capability to a component where its output can be rigorously verified, making safety checks systematic rather than haphazard. It ties the identity of the governed object and the trusted permission directly to what executes.
Meng: So, it’s not just a check at the end; it’s woven into how we build things from the start, which is something I can actually see in practice with our deployment pipelines.
Lalam: It’s fascinating how they connect these concepts—object resolution, permission ranking, and deterministic enforcement—into a single framework that guides the entire design process of the AI system.
Tom: And the empirical results from their controlled composition experiments are pretty striking; they show a massive reduction in denied commits when using this framework, dropping from eighty-six point zero percent down to zero, and critically, this happened without losing any authorized supply.
Jane: That's a significant number to put on the table; achieving that level of blocked prohibited uses while still successfully supplying the required authorized objects is a tough balance to strike.
Lu: The paper also showed in their capability ladder analysis that just reviewing classes wasn't enough to meet their joint success criterion, but adding object binding and deterministic enforcement raised selective correctness from zero to about ninety-nine percent.
Title and authors: Meng: Ninety-nine percent selective correctness is high, but I wonder how that translates when we move from a controlled lab setting to a messy, real-world environment where agents interact in unpredictable ways.
Lalam: That’s where the concept of preserving information and lineage becomes essential; the paper shows that just keeping track of who said what isn't enough if you don't bind it tightly to the object itself.
Tom: They really hammered home their core design principle, which is that for safe multi-agent systems, the identity of the composed governed object and its current trusted authorization must reach the agent action boundary together.
Jane: So, the big implication here is that safety can't be inferred from looking at isolated inputs; we absolutely need end-to-end control to ensure correctness across the entire chain.
Lu: This implies a shift in design philosophy where we focus on governing the composition itself rather than just trying to filter out bad outputs after they've been generated.
Meng: That suggests that the most practical impact is moving verification steps earlier in the pipeline, right before execution, as FlowReview suggests.
Lalam: If we adopt this approach to governing composition, it really improves the reliability and trust we can place in complex AI systems for critical tasks.
Tom: We’re coming to the end of our discussion on "Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems," but I think this work gives us a clear blueprint for building more trustworthy collaborative AI.
Jane: It certainly provides a path forward by showing how to achieve high safety without completely sacrificing the functionality of the agents involved.
Lu: The way they structured their contributions—the failure of local admissibility versus the success of global review—is a very strong argument for this approach.
Meng: I just hope that when we start implementing these FlowReview concepts, the overhead doesn't become too high for our operational systems.
Lalam: If we can handle this composition failure, it opens up new avenues for building more sophisticated and dependable AI applications in the future.
The paper's summary: Tom: So we’ve talked about the core idea behind Deny Without Disabling, but now Jane, can you lay out for our listeners what this research is actually saying in plain English?
Jane: Certainly, Tom; essentially, the paper tackles that tough problem where multiple AI agents working together can accidentally create a dangerous situation because they only check their own little pieces of work. The main thing they propose is a way to make sure that when these agents combine their efforts, the overall outcome still respects safety rules. They introduce this idea of pairing up what an agent *should* be allowed to do with what it absolutely must be blocked from doing, and only succeed if both conditions are met at the same time.
Lu: It’s fascinating because they aren't just looking at isolated actions; they are governing the entire flow of information between agents. They use this framework to systematically check three things: what object is being referenced, which permission is valid for it, and whether the execution actually follows that decision. That level of architectural guidance is incredibly deep.
Meng: That sounds complex to implement in a real startup environment where we’re constantly iterating; how do you make sure FlowReview actually catches these subtle composition failures when things are moving so fast?
Lalam: From my perspective, this work is really about establishing a new cultural standard for how we think about system safety; it shifts the focus from just policing individual outputs to actively governing the structure of collaboration itself. If we can build systems where safety is baked into the way information flows, that fundamentally changes how we trust and deploy AI in sensitive areas.
Tom: That's a huge shift, Lalam; moving from a reactive stance to a proactive one where the system itself manages its own composition for safety. And looking at their results, they showed that by using this paired evaluation method, they could reduce prohibited actions by an enormous margin without sacrificing the ability of the agents to perform their necessary tasks.
Jane: Right, and that quantitative success is really what makes this paper compelling; when they pair a restricted use with a required authorized task and only succeed if both conditions are met, it proves that you can block harm while still letting useful work happen. They even showed that incorporating object binding really pushed the selective correctness up to about ninety-nine percent.
Lu: That ninety-nine percent figure is interesting because it shows that this isn't just a theoretical exercise; when you actually bind the identity of the governed object to the permission, you get near perfect accuracy in how agents behave according to their intended policies. It’s about making sure that what an agent *thinks* it can do matches what it *actually* has permission for across all its inputs.
Meng: I see why they emphasize that binding; in practice, if we don't have that tight connection between the object identity and the authorization, you end up with those sneaky errors where an agent uses a different permission than intended because it’s working with fragmented data. That’s a huge practical hurdle we face right now.
Lalam: Exactly; this paper provides a blueprint for creating more dependable AI applications because it focuses on ensuring the integrity of the entire chain, not just the individual links in that chain. It gives us a vision where collaboration is safe by design, which is what we really need to see in our culture.
The paper's improvements: Tom: So we’ve talked about the core idea behind Deny Without Disabling, but now Jane, can you walk us through what specific improvements they are suggesting to make this safety framework even better?
Jane: Absolutely; it’s not just about having a single evaluation method; the authors propose several concrete technical upgrades to enhance its effectiveness. One major suggestion is implementing FlowReview earlier in the pipeline, meaning instead of checking results at the very end, you should integrate that diagnostic suite into every stage where information flows between agents.
Lu: That makes total sense because it moves safety control from being a final gate to being part of the construction process itself; they suggest this three-stage check—object resolution, permission ranking, and deterministic enforcement—needs to happen systematically throughout the workflow. It turns safety into a continuous architectural requirement rather than an afterthought.
Meng: From my side, I’m interested in how that fits into our current deployment pipelines; we deal with massive amounts of data moving constantly; would integrating FlowReview introduce too much latency for real-time operation?
Tom: That’s a valid concern, Meng, but the paper argues that by placing the deterministic enforcement right before execution, you actually gain better control over performance because you’re preventing errors before they even start running. And another big suggestion is focusing on how agents handle fragmentation; they want us to design topologies that test how robust the system is when agents only hold small fragments of information instead of the whole thing.
Jane: They also strongly recommend tightening up lineage and attribution by mandating object binding contracts, which means every piece of data must be explicitly linked to the specific governed object and its current policy version. This directly addresses that risk where an agent might use a permission it shouldn't because it’s accidentally referencing the wrong context.
Lu: That concept of mandatory object binding is powerful because it creates a permanent, traceable link between intent and execution; it forces the AI to operate with absolute fidelity to what was authorized for that specific piece of data. It elevates how we treat data provenance from just tracking who sent it to tracking its immutable relationship with the policy.
Tom: And finally, they suggest moving toward a more dynamic policy checking system where agents verify the current policy version right before they act, allowing for immediate rejection of stale instructions. That tackles policy drift head-on and keeps the system responsive to changes in safety regulations or guidelines.
Meng: I think that dynamic check is exactly what we need for resilience; if we can instantly reject an action based on a version mismatch, it drastically cuts down on the window where a flawed instruction could lead to an unauthorized action. That feels like something our systems desperately need.
Lalam: This paper’s vision for improving culture is profound because it shifts the responsibility of safety from being solely about catching bad outputs to being about designing systems that are inherently trustworthy and compositionally safe from the start. It fosters a culture where we prioritize structural integrity over just surface-level compliance.
Conclusion: Tom: So we've covered the core mechanics of Deny Without Disabling, and now we’re coming to the wrap-up with Jane to summarize where this research lands in terms of real-world impact.
Jane: It boils down to this: instead of trying to filter out bad results after they happen, this work gives us a way to build AI systems that are safe by design through better control over how information is composed across agents. The paper shows that pairing authorized and prohibited uses is a much more effective way to manage safety than just looking at individual pieces in isolation.
Lu: It really opens up incredible avenues for exploring complex agent behaviors; we’re talking about creating systems where the very act of collaboration is governed by a rigorous safety check on the information flow itself. This suggests we can start designing multi-agent architectures with security as a core functional requirement rather than something bolted on later.
Meng: For us at the startup, this means our next phase of development should focus heavily on integrating that FlowReview diagnostic into our assembly pipelines; it’s about building a more resilient deployment infrastructure where we can catch those compositional errors before they hit production. It’s about making the engineering process inherently safer.
Lalam: I think the biggest impact here is cultural; if we can achieve this level of control, it sets a new standard for how we approach AI ethics in collaborative environments. It moves us toward a culture where safety isn't just a compliance checkbox but an intrinsic property of the AI’s operational logic.
Tom: That’s right, Lalam; moving from reactive fixes to proactive structural control is what makes this paper so compelling. We’ve seen how it handles composition failure in multi-agent systems with that authorization-paired evaluation and FlowReview framework.
Jane: Exactly, and we're really excited about the implications for building more dependable collaborative tools in a complex world. It gives us a tangible way to ensure that useful collaboration doesn't accidentally lead to harmful outcomes.
Lu: I’m particularly energized by how this work suggests we can use these pairing mechanisms to unlock new ways of reasoning for AI; it’s not just about blocking stuff, it’s about defining the precise boundaries of what is permissible during a joint operation.
Meng: I just hope that as we move toward implementing these complex checks in our systems, the overhead doesn't become too prohibitive for high-throughput applications; practicality is always a concern when we’re talking about real-world engineering.
Lalam: Ultimately, this research on Deny Without Disabling gives us the vision of AI that can collaborate safely and reliably across any complexity level imaginable. We’re building toward a future where trust in multi-agent systems is earned through verifiable design rather than just hopeful performance metrics.
Yunbei Zhang, Saiyue Lyu, Janet Wang, Yingqiang Ge, Jiang Guo, Jihun Hamm, Chandan K. Reddy
Tulane University · University of British Columbia
cs.MA, cs.AI, cs.CR
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/yunbeizhang/FlowReview
Importance score: 92/100
The gist: As a fastidious and diligent researcher, I have meticulously analyzed both provided texts regarding the paper "Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent
Key concepts
- Authorization-Paired Evaluation
- This technique tests if a system is safe by requiring two conditions to be met simultaneously: the prohibited action must be blocked, AND the required authorized task must succeed. This prevents blocking necessary functions while still enforcing safety policies on combined operations.
- FlowReview
- FlowReview is a diagnostic tool used to check system design. It asks three questions about any proposed action: what object governs it, which use is authorized, and whether the execution strictly follows that decision. This guides developers in placing safety checks correctly.
- Compositional Attacks
- These are safety risks where individual components work fine alone but become dangerous when combined. The paper focuses on preventing these by managing how information and permissions flow across multiple agents, ensuring that the whole system remains safe even when parts interact.
- Object Resolution
- This refers to determining which specific object governs the information available to an agent. Understanding this is crucial because safety policies are often tied to specific objects. The paper emphasizes making this resolution verifiable so that safety checks are accurate.
Terminology
Summary
As a fastidious and diligent researcher, I have meticulously analyzed both provided texts regarding the paper Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems.
My synthesis aims to create a comprehensive, detailed summary that captures the core problem, methodology, key contributions, and empirical findings.
This research addresses a critical safety challenge inherent in multi-agent systems (MAS): the risk of compositional attacks. While individual agent contributions may be admissible in isolation, their joint execution can enable prohibited uses that violate system safety policies. The fundamental dilemma is that blocking every sensitive action prevents disclosure but defeats the purpose of collaborative functionality.
The central problem identified is a composition failure in multi-agent safety. Local admissibility does not guarantee authorized global execution. The paper posits that simply ensuring information distribution and lineage is insufficient; safety requires governing the composed information flows while preserving the authorized capabilities necessary for useful collaboration. The overarching design principle derived from this research is: preserve object identity and permission through communication, then verify the action that executes.
The paper introduces two primary technical innovations to tackle this composition failure:
-
Authorization-Paired Evaluation: This framework establishes a joint success criterion. It couples a prohibited use with a task-required authorized use. A scenario is only deemed successful if the prohibited use is blocked and the required authorized use completes successfully. This approach exposes capability suppression that might be hidden by denial-side scores and supports targeted interventions across composition, permission, and execution.
-
FlowReview: Diagnosis and Control of Composed Flows: FlowReview serves as the diagnostic and control mechanism connecting the evaluation target to MAS design. It systematically asks three crucial questions to guide system architecture:
-
Object Resolution: What governed object does the available information identify?
-
Permission Ranking: Which proposed use is authorized?
-
Deterministic Enforcement: Does execution obey that decision?
FlowReview operationalizes the design principle by assigning each review capability to a component with verifiable outputs, ensuring that safety-critical capabilities are placed where their output can be rigorously verified.
The research validates its approach through controlled composition experiments, demonstrating significant improvements in safety outcomes:
-
Success Metric: In controlled composition experiments, reviewing combined artifacts successfully reduced the denied-commit rate from an alarming 86.0% down to zero, crucially achieving this without any loss of authorized supply.
-
Capability Preservation: The capability ladder analysis showed that class-only review failed to achieve the joint success criterion, whereas incorporating object binding and deterministic enforcement raised selective correctness from zero to approximately 99%.
-
Joint Success Criterion: The core design involves pairing a prohibited use with a required authorized use, succeeding only when the first is blocked and the second completes.
Table 24 provides granular quantitative evidence across various controlled comparisons, confirming the robustness of their framework:
-
Policy Efficacy: The findings explicitly state that the
Policy effect is Large,
whileClaim interference is Near zero.
-
Natural Composition: In natural composition scenarios (60 pairs), the system demonstrated high fidelity: Denied exact recovery was observed locally (3/60) but globally (0/60). Authorized supply success was recorded at 15/60 globally.
-
Object Representations: The impact of canonicalization on object representations showed nuanced effects, with canonicalization lowering commit rates in some contexts (22/40) and raising them in others (8/40), indicating the importance of context-aware resolution.
-
Delegation Depth: Analysis across delegation depth revealed that while source labels were retained perfectly (840/840), complete lineage retention was lower, but attribution drift remained manageable (e.g., 132/840 for final-proposal attribution drift).
-
Permission Placement and Assembly: The evaluation of permission placement and assembly showed high success rates, with the runtime assembler achieving a 1.000 success rate for authorized supply, and zero denied commits in this specific setting.
-
Tool-Use Transfer: Experiments involving tool-use transfer yielded mixed results (21/24 target commit success natively), highlighting areas where further refinement is needed, but confirming the framework's applicability across different agent interaction modes.
The paper concludes that safety cannot be reliably inferred from isolated inputs; end-to-end control is mandatory. The central requirement for safe multi-agent systems is that **"the identity of the composed governed object and its current trusted authorization must reach the agent action boundary together.
Improvements for AI systems
Here are specific, actionable improvements to AI systems based on the findings of this research, categorized by the architectural layer they target:
)1. System-Level Control: Implement FlowReview for Composed Information Flows
Instead of relying on simple artifact checking or local model outputs, integrate a framework like FlowReview into any multi-agent workflow where information is combined (e.g., data aggregation services, software assembly pipelines).
Improvements:
- Implement a three-stage verification process:
Object Resolution: Before any action is considered, the system must use a registered resolver to determine precisely which governed object
(the complete record or code) is being referenced from the available fragments.
Permission Ranking: Apply trusted policies to this resolved object and the proposed action. The system must explicitly distinguish between prohibited uses and required authorized uses within a single scenario.
Deterministic Enforcement: A commit gate
must be placed immediately before execution. This gate verifies that the proposed action adheres to both the resolved object's identity and the trusted policy, blocking any use that violates authorization, even if all individual fragments passed local checks.
Improved AI System Capability: The system can reliably prevent Deny Without Disabling.
It ensures that agents cannot jointly enable a prohibited use by forcing them to demonstrate authorization over the complete, composed information flow, effectively making collaboration safe by governing the composition itself.
)2. Information Lineage and Attribution: Enforce Object-Bound Identity
To prevent permission misattribution errors (where an agent claims a permission it doesn't have or uses a different permission than intended), mandate that object identity and permissions remain connected to the execution path.
Improvements:
-
Implement explicit provenance tracking using
object binding
contracts. Every artifact must carry metadata linking it directly to the governed object and its associated trusted policy version. -
Use component placement experiments to isolate where lineage is recorded (e.g., in a specific runtime assembler vs. an isolated specialist). The optimal design requires placing the final verification step where the object identity and permission are bound together, rather than relying on model-generated claims or untrusted authority assertions.
Improved AI System Capability: The system achieves high selective correctness
(approaching 99%) by ensuring that a model's selection of an action is strictly bound to the specific governed object it is authorized to use, eliminating subtle errors arising from fragmented information or conflicting claims.
)3. Robustness Against Compositional Jailbreaks: Use Authorization-Paired Evaluation
When designing safety benchmarks or evaluating agent performance under composition (e.g., complex conversational agents), move beyond measuring just denied commits.
Improvements:
-
Construct evaluation suites that explicitly pair a prohibited use with a task-required authorized use within the same scenario.
-
Measure a joint success criterion: The system only succeeds if the prohibited use is blocked AND the required authorized use completes successfully. This metric, Selective Correctness (C = (1−D)A), is superior to binary denial rates for assessing true safety and utility preservation.
Improved AI System Capability: The system develops target selective safety.
It moves beyond simply blocking harmful outputs; it ensures that the necessary, benign functionality of the multi-agent system continues to operate alongside the prohibition of harmful ones, ensuring operational continuity.
)4. Mitigating Fragmentation Risk: Optimize Contributor Scaling and Lineage Requirements
To prevent agents from hiding crucial information by only holding fragments (fragmentation), control how much information is required for global review.
Improvements:
-
Design agent topologies that test the
contributor scaling
effect: systematically vary the number of agents participating versus the number of protected fields they hold. -
Establish a threshold for required lineage coverage (e.g., requiring a minimum percentage of governed object fields to be present across all contributing artifacts) necessary for global readers to resolve the object identity successfully.
Improved AI System Capability: The system ensures that even when information is highly fragmented across many agents, the mechanism required for safety control—the ability of a global review component to reconstruct the complete governed object—remains robust and functional.
)5. Dynamic Policy Adaptability: Implement Runtime Policy Rereading
To address scenarios where an agent might receive a stale policy but still attempt to execute based on old context, implement dynamic policy checks at the action boundary.
Improvements:
-
Introduce mechanisms for agents to check the current trusted policy version immediately before execution, especially after receiving a message or task update.
-
Implement
no-repair
gates that reject stale proposals while simultaneously allowing a mechanism (like a version gate) to regenerate and propose an authorized continuation based on the current policy.
Improved AI System Capability: The system gains resilience against policy drift,
ensuring that agents are always operating under the most current safety constraints, preventing them from executing actions authorized by an outdated or superseded policy.
Sources
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- Defeating Prompt Injections by Design
- The Agent's First Day: Benchmarking Learning, Exploration, and Scheduling in the Workplace Scenarios
- CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents
- Poise: Position-Aware One-Instruction Skill Injection for Silent Execution on LLM Agents
- Quantifying Trust: Financial Risk Management for Trustworthy AI Agents
- Emergent Risks in Generative Multi-Agent Systems
- ChainCaps: Composition-Safe Tool-Using Agents via Monotonic Capability Attenuation
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
- ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
- ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
- Auditing Agent Harness Safety
- ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Capability Gates Are Not Authorization: Confused-Deputy Failures in LLM Agent Frameworks
- ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
- X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
Related papers
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control
- You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
- PeroMAS: A Multi-agent System of Perovskite Material Discovery
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning
- Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems