Capability Sheaves for Compositional Agent-Harness Repair: Controlled Quotients and a Real-Repository Stress Test

arXiv:2608.13228 · cs.AI, cs.LG · Submitted 2026-08-13 · Read on arXiv

Saveliy Batruin

cs.AI, cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 72/100

The gist: The paper introduces a finite capability sheaf to model failures in AI agent harnesses, where locally successful components disagree on shared state.

Terminology

Summary

The paper introduces a finite capability sheaf to model failures in AI agent harnesses, where locally successful components disagree on shared state. The sheaf has typed behavior stalks for five requirements (localization, contract, ordering, preservation, verification), restriction maps that are literal field projections, and six overlaps. An exact finite constraint-satisfaction problem (CSP) defines acceptance, while a linearized relative cohomology class provides a diagnostic and search feature.

The paper proves several theoretical results: Theorem 1 (Augmented feasibility) states that a candidate is a global section accepted by the cover-local predicate if and only if every local predicate holds and every registered pair of restrictions agrees. Proposition 2 shows that a nonzero quotient class is a sound no-candidate certificate, but the converse is false. Theorem 3 (Local-score repair-order separation) proves that a top-one selector using only isolated local signatures has minimax success at most 1/K, while a zero-class-first selector ranks the feasible candidate first. Theorem 4 (Rank-truncated residual stability) bounds the distance between estimated and population residuals under perturbation, and Corollary 5 gives a finite-trace guarantee. Theorem 6 (Finite-trace repair recovery) shows that with enough fresh traces, the empirical maximizer of the weakest-face score recovers the true candidate.

In a controlled experiment over 20 task clusters with hidden interior mediators, quotienting their coboundaries reduces the candidate budget from 2,000 to 1,000 per cluster. The aligned-state ablation removes the gap, and exact CSP matches the quotient. The exact one-sided sign-test value is p = 9.5367 × 10−7, with all 20 clusters favoring the quotient. This demonstrates invariance to stale representatives, not superiority over exact reasoning.

The real-repository stress test uses a discovery split from the SWE-bench Multilingual pool of PatchFuseBench: 160 issues from 20 repositories, 875 real candidate patches, 2,579 source-aware edit atoms, and 153 newly executed patches. A first pool-level construction is constant because [b − Dx] = [b] in coker D and therefore cannot rank configurations. A candidate-indexed repair is nontrivial on 848/875 candidates and varies within 120/160 issues. It resolves 118 issues versus 116 for a matched noncohomological selector, but the difference is not supported across repositories (exact sign-flip p = 0.75). A leave-one-repository-out abstention gate reaches 127/160, tying the strong anchor and exceeding its matched gate by one issue (p = 1.0). The discovery gate therefore fails and the confirmatory split remains sealed.

The paper concludes that the controlled invariance mechanism and an identifiability correction are supported, but not a real-world cohomological advantage. The main scientific result is a boundary: hidden-state quotienting succeeds in the controlled intervention, while the tested obligation-and-risk complex is too weak for real patch fusion.

Improvements for AI systems

Improvements to AI Systems:

  1. Add a cohomological consistency layer to multi-agent harnesses.
  • The AI system can detect and diagnose disagreements between locally successful components (e.g., planners, executors, verifiers) on shared state by computing a relative cohomology class.

  • It can flag candidates whose coboundary is nonzero as no-certificate without exhaustive search, reducing candidate budgets (e.g., from 2,000 to 1,000 per cluster) and pruning invalid configurations early.

  1. Implement a zero-class-first ranking selector for repair search.
  • Instead of ranking by isolated local scores (which has minimax success ≤ 1/K), the system can prioritize candidates with zero quotient class (i.e., globally consistent) before any local-score-based ranking.

  • This guarantees the feasible candidate is ranked first, improving repair recovery in finite-trace settings and reducing wasted compute on inconsistent states.

  1. Use rank-truncated residual stability for robust evaluation under perturbation.
  • The system can bound the distance between estimated and population residuals, enabling reliable early stopping or trace-budget allocation.

  • It can decide when to stop collecting fresh traces based on a finite-trace guarantee, avoiding overfitting to noisy local signals.

  1. Add an abstention gate based on leave-one-repository-out validation.
  • The system can refuse to make a repair recommendation when the cohomological signal is not stable across subpopulations (e.g., repositories), preventing overconfident deployment on unseen domains.

  • This improves reliability in real-world patch fusion, where a naive cohomological selector fails to outperform a matched baseline.

  1. Integrate a finite constraint-satisfaction problem (CSP) acceptance check.
  • The system can use an exact CSP to verify global sections (candidates) against cover-local predicates and pairwise restrictions, ensuring that only fully consistent candidates are accepted.

  • This provides a sound and complete acceptance mechanism, matching the quotient-based diagnostic without false positives.

What the improved AI system can do:

  • Reduce search space and compute by pruning inconsistent candidates via cohomology.

  • Rank feasible solutions first, improving repair success with fewer traces.

  • Provide robust confidence bounds under noisy or perturbed inputs.

  • Abstain from acting when the underlying model is too weak or domain-shifted.

  • Guarantee logical consistency of multi-agent outputs on shared state, preventing silent failures from local successes.

Abstract

Agent harnesses combine retrieval, routing, state, provenance, and verification, but locally successful components may disagree on shared state. We model this failure with a finite capability sheaf: stalks encode typed behavior signatures, restriction maps retain shared fields, and accepted runs are useful global sections. An exact finite constraint-satisfaction problem (CSP) defines acceptance, while a linearized relative cohomology class provides a diagnostic and search feature. A controlled experiment over 20 task clusters introduces hidden interior mediators whose raw states are nuisance variables. Quotienting their coboundaries reduces the candidate budget from 2,000 to 1,000 per cluster; aligning the hidden state removes the gap. Exact CSP matches the quotient, so the result demonstrates invariance to stale representatives, not superiority over exact reasoning. We then test the method on a discovery split from the SWE-bench Multilingual pool of PatchFuseBench: 160 issues from 20 repositories, 875 real candidate patches, 2,579 source-aware edit atoms, and 153 newly executed patches. A first pool-level construction is constant because [b-Dx]=[b] in cokerD and therefore cannot rank configurations. A candidate-indexed repair is nontrivial on 848/875 candidates and varies within 120/160 issues. It resolves 118 issues versus 116 for a matched noncohomological selector, but the difference is not supported across repositories (exact sign-flip p=0.75). A leave-one-repository-out abstention gate reaches 127/160, tying the strong anchor and exceeding its matched gate by one issue (p=1.0). The discovery gate therefore fails and the confirmatory split remains sealed. The study supports the controlled invariance mechanism and an identifiability correction, but not a real-world cohomological advantage.

Sources

Related papers