A Verifier Can Leak the Answer: Diagnosability Before Optimization in Closed-Loop Agent Debugging

arXiv:2610.00126 · cs.CR, cs.AI, cs.MA · Submitted 2026-09-09 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "A Verifier Can Leak the Answer".

Elias: A verifier-grounded evaluation shows that an optimizer can appear effective without resolving genuine ambiguity if its probes or predicates encode the target identity.

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So we're diving into "A Verifier Can Leak the Answer: Diagnosability Before Optimization in Closed-Loop Agent Debugging" today. This paper tackles a really subtle problem where an optimizer can look good even when it isn't actually solving anything meaningful because of how the verifier is set up.

Elias: Exactly, Nadia. The core thesis here is that if the verifier's probes or predicates accidentally encode what the target identity is, then you can compare different solvers and find they both look equally effective without actually finding a real fault or ambiguity to resolve.

Priya: From my side, I'm curious about what this means for the actual data we collect; does this leakage impact how accurately we measure the privacy or detection power of these agent components?

Nadia: That's a fair question, Priya. The paper points out that in their aggregate-trace debugger for a closed-loop decision agent, an exact minimum hitting set optimizer and a propagation-aware greedy method returned identical supports in twelve out of twelve development cases.

Elias: That's the vacuous comparison they're highlighting; the two solvers were basically giving the same results because of how the evidence was structured within that specific setup.

Priya: So, if we look at what this means for measurement, does it suggest that just having a larger set of components or more traffic exposure isn't enough to guarantee we're getting meaningful diagnostic information?

Nadia: Precisely. The paper identifies a failure mechanism where an exact-component predicate and a one-component hard probe interacted to create planted component singletons, which then got propagated, meaning neither solver had a real choice afterward.

Elias: That interaction is key; the authors found that the compiled conflict family was already solved in those development results because of how the evidence was constructed.

Priya: So, if we think about real-world measurements, this suggests that when we test a system, we need to be careful that our testing setup isn't accidentally telling the agent what it's going to do before it gets a chance to make a real decision.

Nadia: Exactly. To fix this leakage, they propose introducing a two-stage diagnosability gate instead of just running solver evaluation first.

Elias: That gating mechanism involves a clean reference-map gate and then a matched runtime two-stream gate, requiring specific support counts in those partitions for admission to continue.

Priya: What does that admission process actually look like from the perspective of privacy or measurement researchers? Are we talking about setting minimum thresholds on how much traffic or evidence needs to be present before we trust the results?

Nadia: It's more than just a threshold; they independently calibrated stable false admission at the physical-component level, separate from coverage and detection power. They also showed that in admitted cases, affected clean traffic is a better power coordinate than structural fault size.

Elias: That distinction between traffic exposure and structural fault size is important; it suggests that the type of observation matters more than just how big the component or the fault itself is.

Priya: So, if this holds up in their heldout experiment, does it give us a better idea about what kind of evidence we should be prioritizing when trying to assess agent safety or privacy guarantees?

Nadia: The formal heldout experiment showed the gated verifier was operational, admitting fifty-five out of seventy-two units at reference and rejecting one additional represented component at runtime.

Elias: And they confirmed there were no stable false admissions among the twenty represented components in that test, which suggests the new structure is working to prevent that leakage.

Priya: That's reassuring for anyone looking at agent debugging; it sounds like a structural change to the verification process rather than just tweaking an algorithm.

Nadia: It really is, and the authors laid out some critical structural lessons for future agent development, like evidence eligibility and non-revelation need to be verified before optimizing component selection.

Elias: I agree with that point about verification coming first; a stronger solver can only certify a stronger verifier artifact, which is what the paper cautions against.

Priya: So, if we look at the broader implication for the world, this suggests that in complex agent systems where decisions are made in closed loops, the way we verify those decisions is just as important as how smart our optimization algorithm is.

Nadia: Right. The final result they present is a modular claim ledger, which separates different claims like execution integrity or low false admission into distinct endpoints.

Elias: That separation prevents selective interpretation of the results, ensuring that an agent update can improve traffic exposure without compromising diagnosis quality.

Priya: For the folks in privacy research, this means we have a clearer path to ensure that safety and diagnostic quality are maintained even when we're trying to optimize for different things, like coverage versus detection power.

Nadia: So, looking at the title of "A Verifier Can Leak the Answer: Diagnosability Before Optimization in Closed-Loop Agent Debugging," it really summarizes this tension between verification and optimization.

Elias: It highlights that we need to establish evidence eligibility and non-revelation before we even think about optimizing the component selection, because otherwise, the comparison is meaningless.

Priya: I think it means that for any system we build involving agents making decisions in a loop, the foundation of how we verify those decisions has to be solid before we start chasing optimization improvements.

Conclusion: Nadia: So, we're wrapping up our discussion on "A Verifier Can Leak the Answer: Diagnosability Before Optimization in Closed-Loop Agent Debugging," and I want to quickly recap how this paper shows that if the verifier is set up wrong, even a sophisticated optimizer can be misleading.

Elias: It really boils down to how the probes or predicates in the verifier can accidentally reveal the target identity, making any comparison between different solvers essentially meaningless if you haven't verified that leakage is absent.

Priya: And from my side, I'm still focused on what this means practically for the data we collect; does this structural issue translate into a problem with how accurately we measure privacy or detection power in these closed-loop systems?

Nadia: Exactly, Priya, because the authors show that the two solvers agreed on results because of a specific interaction between an exact predicate and a probe that created planted singletons, which was already solved before optimization could really decide anything new.

Elias: That interaction is what's worrying from a cryptographic standpoint; it suggests that if we rely on solver comparisons without verifying this structural integrity, we might be trusting results based on a flawed premise about the evidence eligibility.

Priya: So, if this leakage happens in real-world testing, it could mean that our measurements of system safety or privacy are actually being skewed because the verifier was already giving away too much information before we even started optimizing.

Nadia: Precisely; the authors propose a two-stage diagnosability gate to stop this leakage by establishing evidence eligibility and non-revelation before any component selection optimization takes place.

Elias: That gating mechanism sounds like a necessary safeguard, but I wonder if the complexity of setting up that gate itself introduces new vulnerabilities that we might not have accounted for in the initial proof assumptions.

Priya: It seems like a necessary step to ensure our data reflects actual system behavior rather than artifacts of an overly permissive verification setup, which is crucial for any serious privacy research.

Nadia: Absolutely, and the paper's conclusion points toward a modular claim ledger to keep different metrics—like execution integrity versus low false admission—separate so we don't get selective interpretations later.

Elias: That separation sounds like a smart way to manage complexity, but it doesn't solve the fundamental issue if the initial evidence being fed into that ledger is already compromised by a leaky verifier.

Priya: So, the core message is that reliable agent development requires rigorously verifying the evidence generator itself before we ever trust an algorithm designed to optimize its output, which sets a high bar for our measurement efforts moving forward.

Blossom AI Blossom AI Labs

cs.CR, cs.AI, cs.MA

Submitted: 2026-09-09

Updated: 2026-09-09

Comments: Submitted to Who Verifies the Agents? Toward Reliable Agent Development (NeurIPS 2026 workshop). 7 pages, 0 figures, 2 tables. The reproducibility artifact is linked in the paper

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 90/100

The gist: A verifier-grounded evaluation shows that an optimizer can appear effective without resolving genuine ambiguity if its probes or predicates encode the target identity.

Key concepts

Vacuous Comparison
This occurs when comparing two solvers where the verification process itself encodes the correct answer. If the probes or predicates used by the verifier already contain information about which component is faulty, then any optimization performed on those solvers becomes a comparison of identical outcomes, rendering the solver performance metric useless.
Evidence Eligibility
This refers to whether a piece of evidence is useful for diagnosis. Evidence is eligible if the agent visits components frequently enough that aggregate traces provide meaningful information. It also requires that reference and current observations are comparable, ensuring the evidence can actually be used to distinguish between different possibilities.
Non-Revelation
This concept means the diagnostic intervention must not inadvertently reveal the hidden answer. A verifier is non-revealing if its scope or mechanics do not encode the target identity. If a verifier leaks information about the fault location, it bypasses the need for a complex optimizer to find it.
Support-Gated Verification Contract
This proposed solution replaces solver-first evaluation with a two-stage gate. The first gate checks if existing support is high in clean partitions. The second, matched runtime gate requires both reference and current streams to have high support simultaneously, ensuring the verification process is robust before proceeding to optimization.

Terminology

Summary

A verifier-grounded evaluation shows that an optimizer can appear effective without resolving genuine ambiguity if its probes or predicates encode the target identity. This research demonstrates that in closed-loop agent debugging, a verifier must establish evidence eligibility and non-revelation before optimizing component selection.

The core problem addressed is the vacuous comparison between solvers.

The paper investigates a closed-loop decision agent to study how aggregate traces and diagnostic pipelines determine fault localization. The central issue examined is whether an exact Minimum Hitting Set (MHS) optimizer provides a substantive advantage over a propagation-aware greedy method when the verifier's probes inherently encode the target identity, rendering the optimization comparison meaningless. The finding is that Exact MHS and a propagation-aware greedy method returned identical supports in 12/12 development cases and the same planted-fault recovery in 9/12.

The failure mechanism is identified through an audit of development results.

The audit revealed that the two solvers agreed because the compiled conflict family was already solved. This occurred because an exact-component predicate and a one-component hard probe interacted to create planted component singletons. Once singleton propagation ran, neither solver had a substantive choice. The paper explicitly states, This is a verifier failure, not a mathematical failure of MHS, suggesting that the verifier itself disclosed the answer before optimization could resolve genuine ambiguity.

The proposed solution introduces a two-stage diagnosability gate.

To prevent this leakage, the authors propose replacing solver-first evaluation with a support-gated verification contract. This contract involves a sequence of checks:

  1. A clean reference map gate, where support is established based on whether its existing support count is at least 12 in clean partitions.

  2. A matched runtime two-stream gate, which requires that both its reference stream and its current stream have support of at least 12 in that same partition. Admission requires at least 14 jointly supported partitions out of 15.

The formal heldout experiment validates the new verifier structure.

In a preregistered heldout comprising 1,440 cases, the gated verifier was tested. The results confirmed its operational status: it admits 55/72 units at reference, rejects one additional represented component at runtime, and produced no stable false admissions among 20 represented components. Furthermore, the analysis showed that affected clean traffic predicted detection better than nominal fault-cell fraction, indicating that traffic is a more relevant power coordinate than structural fault size.

Key structural lessons guide future agent development.

The paper outlines four critical conditions for a verifier to support an optimizer comparison:

  1. Evidence must be eligible, meaning the policy visited components often enough for aggregate traces to carry information, and reference and current observations are comparable.

  2. The evaluation intervention must be non-revealing, meaning it does not encode the hidden answer through its scope or mechanics.

  3. After deterministic propagation, a nontrivial residual decision must remain; otherwise, the solver benchmark cannot support claims about optimizer quality.

  4. Developers should verify evidence eligibility and non-revelation before optimizing the component selector, as a stronger solver can merely certify a stronger verifier artifact.

The final result is a modular claim ledger.

The study concludes by separating different claims—execution integrity, low false admission, adequate coverage, monotone power, and solver superiority—into distinct endpoints. This prevents selective interpretation by ensuring that an agent update can preserve verifier safety while reducing coverage or improving traffic exposure without compromising diagnosis quality. The main takeaway is that Reliable agent development requires verifying the evidence generator before trusting the algorithm that optimizes its output.

The gist: An exact MHS optimizer can appear effective without resolving genuine ambiguity if its probes or predicates encode the target identity. This research demonstrates that in closed-loop agent debugging, a verifier must establish evidence eligibility and non-revelation before optimizing component selection. The core problem addressed is the vacuous comparison between solvers. The failure mechanism is identified through an audit of development results. The proposed solution introduces a two-stage diagnosability gate. The formal heldout experiment validates the new verifier structure. Key structural lessons guide future agent development. The final result is a modular claim ledger.

Table 2: Compact formal verifier outcome

Layer Endpoint Formal result Interpretation

execution protocol, hashes, schedule, validator

56/56 checks; 1,440 cases EXECUTED

evidence reference admission 55/72 units; 20/24 components unsupported units abstain

evidence runtime admission 54/55 reference-admitted units one additional runtime abstention

safety component stable false admission 0/20; upper 0.1391 SAFETY CONFIRMED under 0.20 rule

structure traffic versus cell fraction 0.

Improvements for AI systems

Here are the specific improvements for AI systems based on the findings in this research, structured by architectural change and resulting capability:


  1. A fundamental shift from solver-first evaluation to a support-gated verification contract.

  2. Implementation of a two-stage diagnosability gate: first, a clean reference map gate based on aggregate evidence eligibility; second, a matched runtime two-stream gate requiring comparable support across reference and current observation partitions.

  3. Integration of an independently calibrated false admission threshold (e.g., the 0.20 upper bound) applied at the physical-component level, separate from general coverage metrics.

  4. A modular claim ledger system to prevent score collapse, ensuring that improvements in one metric (e.g., coverage) do not mask regressions in others (e.g., diagnosis quality or safety).

The improved AI system will be able to:

  1. Identify and report when its diagnostic tools are misleading, even if the underlying optimization algorithm is mathematically sound.

  2. Distinguish between missing data (lack of agent exploration/evidence generation) and ambiguity (a genuine, unresolved conflict family).

  3. Automatically abstain from providing a diagnostic score or success label when the evidence gathered by the policy does not meet a pre-defined threshold of eligibility and non-revelation, preventing the system from being misled by weak or structurally flawed probes.

  4. Guarantee that any reported performance metric (like detection power) is grounded in evidence that has been demonstrably generated by the agent's actual execution within comparable environmental conditions, rather than just theoretical potential.

  5. Be rigorously audited to ensure its verification mechanisms do not inadvertently encode the target identity or create artificial singletons that force a premature conclusion on fault localization.

Related papers