Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence
Varun Pratap Bhardwaj, Garima Singh, Arun Pratap Bhardwaj
Qualixar · Independent Researcher
cs.AI, cs.MA
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 49 pages, 12 tables, 25 numbered definitions, 18 theorems with full proofs, six experiments, 65 references. Code, analysis scripts, and preregistration: https://github.com/qualixar/agentassert-abc
Code: https://github.com/qualixar/agentassert-abc
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper investigates the validity of the conditional-independence assumption (C5) that underpins compositional reliability bounds for multi-agent systems, which typically multiply component
Terminology
Summary
This paper investigates the validity of the conditional-independence assumption (C5) that underpins compositional reliability bounds for multi-agent systems, which typically multiply component reliabilities to estimate the reliability of the whole. The authors test this assumption and find it is violated in practice, particularly when components share the same underlying model. They then propose a new, assumption-free method for certifying the reliability of composed agent pipelines.
The core empirical finding comes from a preregistered confirmatory evaluation of 18,000 two-agent handoff missions. The results show that Two instances of mistral-small-24b co-fail on 90.0% of the missions on which either fails, with log OR = 6.66 (95% CI [6.38, 7.00]).
This demonstrates that Two instances of one model do not have independent blind spots. They share them.
The study also shows that substituting a different model reduces this association, but substituting a different vendor (with the model already different) does not, a registered hypothesis that failed to replicate.
The paper proves that the error from assuming independence is signed and runs against the operator. Specifically, under positive dependence, joint failure exceeds the independence product, so redundancy is over-credited precisely when the redundant components share a model.
The authors also show that the obvious alternatives to the independence assumption are flawed: the Fréchet–Hoeffding bounds are often vacuous (the certified floor is zero when mean component reliability is below 1 - 1/m), and fitting a parametric dependence model is worse. They prove in Theorem 4.2 that a bootstrap lower bound on the fitted model’s functional loses coverage of the true reliability as n → ∞, because the identification gap is O(1) while the bootstrap haircut is O(n−1/2).
This means more data makes such a certificate worse, not better.
As a solution, the paper introduces a finite-sample certificate that assumes no dependence structure. This is a linear program over the joint, taken over a Bonferroni–Clopper–Pearson box around measured co-execution moments.
This certificate is sound, sharp for the information supplied, and monotone in the moment family. On real four-stage data, enriching from ten moment functionals to fourteen narrows the identified interval by 85.7% and lifts the certified floor from 0.2455 to 0.4116.
They also provide an anytime-valid certificate that holds its empirical type-I error at 0.0471 or below across every admissible betting fraction, recovering the SPRT exactly at the optimal bet.
Finally, the paper demonstrates that common dependence statistics like Jaccard, ϕ, and Kendall’s τa are bounded by the marginals and can significantly reverse an apparent ordering of conditions when the compared agents fail at different rates.
The authors release all contracts, mission generators, scoring code, analysis scripts, and the preregistration, with every reported statistic regenerated by those scripts.
Improvements for AI systems
Improvements to AI systems:
- Add dependence-aware reliability certificates to multi-agent orchestration.
-
Instead of multiplying per-agent success probabilities, AI systems should compute a certified lower bound on joint mission success using the proposed linear-programming certificate over Bonferroni–Clopper–Pearson boxes of co-execution moments.
-
This gives operators a sound, finite-sample guarantee even when agents share the same underlying model, avoiding over-credit from independence assumptions.
- Replace independence-based redundancy checks with co-failure monitoring.
-
AI systems should track pairwise co-failure rates (e.g., log odds ratio) between agents in production.
-
When two agents share the same base model, the system should flag them as non-redundant and either (a) diversify the model family or (b) adjust the reliability estimate downward using the observed dependence.
- Implement an anytime-valid certificate for online mission reliability.
-
Use the paper’s anytime-valid certificate (with empirical type-I error ≤ 0.0471) to continuously update reliability bounds as new co-execution data arrives.
-
This allows the AI system to stop early or re-route missions when the certified floor drops below a threshold, without waiting for a fixed sample size.
- Avoid parametric dependence models for certification.
-
The system should not fit a parametric copula or correlation model to estimate joint failure, because Theorem 4.2 shows the bootstrap lower bound loses coverage as data grows (identification gap O(1) vs. haircut O(n-1/2)).
-
Instead, use the assumption-free LP certificate, which remains sound and sharp regardless of sample size.
- Use moment enrichment to tighten certificates.
-
The system should dynamically select which co-execution moment functionals to measure (e.g., from 10 to 14) to narrow the identified interval.
-
In practice, this means the orchestrator can improve the certified floor (e.g., from 0.2455 to 0.4116) by collecting richer statistics on joint failures, without changing the agents.
- Correct for marginal-rate confounds in dependence metrics.
-
When comparing redundancy across different agent pairs, the system should not rely on raw Jaccard, ϕ, or Kendall’s τa, as these are bounded by marginals and can reverse ordering when failure rates differ.
-
Instead, use a marginal-adjusted dependence measure (e.g., log odds ratio) or the LP certificate’s identified interval to rank agent pairs for redundancy.
- Preregister and audit mission reliability evaluations.
-
AI systems should adopt the paper’s preregistered confirmatory protocol: define hypotheses, mission generators, and scoring code before deployment.
-
This prevents post-hoc selection of dependence statistics and ensures that reported reliability improvements are not artifacts of model sharing.
What the improved AI system can do:
-
Certify a lower bound on the probability that a multi-agent pipeline completes a mission, even when agents are correlated, with finite-sample guarantees and no distributional assumptions.
-
Detect dangerous co-failure patterns in real time (e.g., two instances of the same model failing together on 90% of shared-failure missions) and automatically trigger model diversification or fallback strategies.
-
Update reliability certificates online as new mission data arrives, maintaining type-I error control without requiring a fixed sample size.
-
Avoid false confidence from parametric dependence models, especially in data-rich regimes where such models become worse.
-
Optimize which statistics to collect (moment functionals) to maximize the certified floor with minimal additional instrumentation.
-
Compare agent pairs for redundancy fairly, using marginal-adjusted measures, so that operational decisions (e.g., which two models to deploy) are not misled by base-rate differences.
-
Provide auditable, reproducible reliability claims by integrating preregistration and script-generated statistics into the system’s reporting pipeline.
Abstract
Compositional reliability bounds for multi-agent systems multiply component reliabilities, a step licensed by a conditional-independence assumption that is routinely stated and rarely tested. We test it. Two instances of one model, in a two-agent handoff, co-fail on 90.0% of the missions on which either fails (log OR 6.66, 95% CI [6.38, 7.00]; phi 0.916), in a preregistered evaluation of 18,000 missions scored by deterministic code with no LLM judge. Substituting a different model reduces the association in six of six contrasts; substituting a different vendor, model already different, does not -- a registered hypothesis reported as a null. The error is signed and runs against the operator: positive dependence inflates joint failure above the independence product, so redundancy is over-credited exactly when components share a model. The assumption-free alternative is often vacuous, and fitting a dependence model is worse: we prove a bootstrap bound on a fitted model's functional loses coverage of the truth as n grows, the identification gap being O(1) while the bootstrap haircut is O(n-1/2). More data makes such a certificate worse, with no visible symptom. We give a finite-sample certificate assuming no dependence structure: a linear program over the joint, over a Bonferroni-Clopper-Pearson box around measured co-execution moments. It is sound, sharp for the information supplied, and monotone in the moment family. Enriching ten moment functionals to fourteen narrows the identified interval by 85.7% and lifts the certified floor from 0.2455 to 0.4116. A companion anytime-valid certificate holds type-I error at 0.0471 under optional stopping. Common dependence statistics are marginal-bounded and can reverse an apparent ordering of conditions when the compared agents fail at different rates. Contracts, scoring code, analysis scripts, and the preregistration are released.
Sources
- Constitutional AI: Harmlessness from AI Feedback
- International AI Safety Report 2026
- Can AI Agents Agree?
- Agent Behavioral Contracts: Formal Specification and Runtime Enforcement for Reliable Autonomous AI Agents
- Multi-Agent Risks from Advanced AI
- Counterfactual Graph for Multi-Agent LLM Calibration
- Harnessing Disagreement: Detecting Correlated Agreement Blindness in Multi-Agent Triage
- VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems
- Governed Capability Evolution: Lifecycle-Time Compatibility Checking and Rollback for AI-Component-Based Systems, with a Proof-of-Concept Evaluation on Embodied Agents
- FALAT: Tracing Failures in LLM Agent Trajectories via Dependency-Guided Search
- Rethinking the Reliability of Multi-agent System: A Perspective from Byzantine Fault Tolerance
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection