Local verification cannot detect non-transportability: a cohomological theory of context preservation in agentic reasoning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Local verification cannot detect non-transportability: Cohomological limits of context preservation in agentic reasoning".
Jane: The paper was written by Suyash Mishra from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the arXiv channel, everyone. Today we’re digging into a paper with a title that sounds like it came from a math department and a philosophy department fighting over the same coffee machine — “Local verification cannot detect non-transportability: Cohomological limits of context preservation in agentic reasoning.”
Jane: Tom, I’m so glad you said that, because when I first read “cohomological limits” I nearly closed the tab. But this paper is actually about something really concrete. It’s about AI agents that take a conclusion from one context — say, a biological finding in one cell type — and carry it over to another context, like a different disease setting. The authors are asking: when can you trust that transfer?
Tom: And their answer is pretty uncomfortable. They prove that the standard way we check these agents — verifying each step locally, making sure the context matches at every bridge — structurally cannot catch a certain kind of error. It’s not that the checks are poorly implemented. It’s that the math says they’re blind to something.
Jane: Right. And the something is called the harmonic component. Let me try to explain it without the topology. Imagine you have three contexts, A, B, and C. You measure the difference between A and B, then B and C, then C and A. If those three measurements don’t add up to zero around the loop, you have an inconsistency. That’s the curl part — and local checks can see it, because you can compare the three pairwise measurements directly.
Tom: But the paper shows there’s another kind of inconsistency that only appears when you go around a bigger loop — through four, five, ten contexts. Every triple of contexts you check locally looks perfectly fine. But when you chain the evidence all the way around, you end up somewhere different from where you started. That’s the harmonic part. And it’s invisible to any check that only looks at small pieces.
Jane: Exactly. And the authors give it a name — holonomy. It’s the same word used in geometry for when you parallel-transport a vector around a curved surface and it comes back rotated. Here, the agent transports a claim around a loop of contexts, and it comes back shifted. The shift isn’t noise. It’s structural.
Tom: So the headline is: an agent can pass every verification check at every step, and still return an answer that’s an artifact of the path it happened to take through the evidence. That’s not a bug in the checks. That’s a property of the context space itself.
Jane: And that’s why the title says local verification cannot detect non-transportability. It’s a proof, not a warning. The authors show that any check supported on a bounded piece of context space — one context, one pair, one triple — cannot distinguish a coherent evidence network from one carrying this hidden loop inconsistency.
Tom: I love that they connect this to something clinicians have known for two decades. In network meta-analysis, they call it loop inconsistency — when direct and indirect evidence around a closed loop of treatment comparisons don’t agree. This paper takes that phenomenon and shows it’s actually the first cohomology of a topological space. It’s the same object, just seen from higher up.
Jane: And that reframing matters, because once you see it as cohomology, you get a toolset. You can measure it, you can decompose it, and you can decide when to abstain. But that’s the next segment — how they actually build a detector and a gating rule out of this. For now, the key point is that the blind spot is real, it’s provable, and it’s not going away with better prompt engineering.
Tom: Stay with us — because the fix they propose is genuinely clever, and it turns the harmonic component from a liability into a data-collection instruction.
Summary: Tom: Back with “Local verification cannot detect non-transportability,” and we’ve established the core problem: local checks can’t see loop-level inconsistencies. But what does the paper actually do about it? Jane, walk us through the method.
Jane: So they build a procedure called Ks.etra — spelled K-s-e-t-r-a, but pronounced like the Sanskrit word for field. The idea is straightforward. You take your context space, you build a network where each context is a node and each overlap between contexts is an edge, and you put the measured evidence on each edge. Then you decompose that evidence into three orthogonal parts.
Tom: And those three parts map onto three different problems. The first part is the gradient — that’s just per-context calibration offsets. If your instrument in one context reads systematically high, that shows up here, and you can fix it by recalibrating. The second part is the curl — that’s the local inconsistency we talked about, the one you can catch by checking triples. And the third part is the harmonic component — the loop-level obstruction that local checks miss.
Jane: Right. And the paper’s central theorem says that only the harmonic component makes your conclusion depend on which reasoning path you took. If the evidence is exact — meaning it comes from a single consistent global assignment — then every path gives the same answer. If there’s harmonic energy, different paths disagree, and that disagreement is not noise. It’s holonomy.
Tom: So what do they do with that? They build an estimator that projects the evidence onto the space of consistent assignments — that removes the gradient and the curl — and then they gate abstention on the harmonic energy alone. If the harmonic component is large relative to your decision tolerance, the agent abstains. And crucially, it doesn’t just say “I don’t know.” It says “I don’t know, and here’s the loop you need to refine to find out.”
Jane: That’s the part I find genuinely beautiful. A non-zero harmonic class isn’t a failure — it’s an instruction. It tells you your context covering is too coarse. There are latent strata inside your nominal contexts that have different true effects, and the only way to resolve the obstruction is to refine the covering — stratify further and collect data there. You can’t fix it by gathering more evidence at the same granularity.
Tom: And they simulate this across two very different domains. One is pharmaceutical real-world evidence, where the context space is contractible — no holes by construction. The other is consumer credit, where the context space includes the business cycle as a circle — expansion, late cycle, contraction, recovery — and that cyclic structure guarantees at least one hole. The credit domain has obstruction baked into the topology of the problem itself.
Jane: The results are striking. They compare their harmonic-gated agent against a baseline that gates on total residual conflict — which is what a careful engineer would naturally build. At matched coverage, the harmonic gate reduces harmful decisions by about three to four percentage points. And the reason is that total conflict is diluted by the gradient and curl components, which carry no information about irreducible error. Only the harmonic part predicts the error that survives optimal estimation.
Tom: And there’s a lovely finance test case in the paper where the null hypothesis is exactly known. In an arbitrage-free market, the log-quote cochain is exactly a coboundary — that’s just a fancy way of saying the exchange rates are consistent with a single underlying value. Triangular arbitrage is the curl. Loop arbitrage is the harmonic. And they show their detector can spot loop arbitrage at one basis point of distortion when a conventional residual monitor needs four.
Jane: So the method isn’t just theoretical. It has a measurable operational advantage. But there’s a twist — the paper includes a correction to one of their own earlier claims, and that’s where we’re heading next.
Improvements: Tom: So we’ve covered the method and the results. But Jane mentioned a correction, and I think that’s actually one of the most honest parts of this paper. What happened?
Jane: So in an earlier version, the authors injected the harmonic component directly into their simulations and then showed it predicted error. That’s circular — of course it predicts error if you put it there. So they redid the whole thing. They removed the injection entirely and let holonomy emerge from a mechanism. And the mechanism is effect modification combined with overlap-specific population composition.
Tom: Meaning the true effect of a treatment differs across latent subgroups, and each overlap between contexts happens to sample those subgroups in different proportions. That’s Simpson’s paradox on a network — every pairwise comparison is internally valid, but they can’t be glued into a global story.
Jane: Exactly. And when they set the effect modification to zero, the harmonic energy drops to machine precision — one point five times ten to the minus fifteen. It vanishes. When they turn it on, holonomy emerges. So the mechanism is real, not injected. But here’s the correction: in the earlier version, they claimed curl energy carries no information about error. That turned out to be an artifact of drawing the curl and harmonic components independently. In the mechanistic model, a single latent cause generates both, so curl is almost as predictive as harmonic.
Tom: So the components are not distinguished by their predictive value — they’re distinguished by their remedy. Curl is repairable by re-measuring the triple where the inconsistency lives. Harmonic is not repairable at that granularity by any amount of data. You have to refine the covering. That’s a genuinely important distinction, and I’m glad they reported the correction rather than burying it.
Jane: They also fixed a second issue. The F-test they derive for the existence of a global claim assumes equal precision across all bridging estimates. In practice, some comparisons are measured much more precisely than others. Ignoring that makes the test anti-conservative — it rejects the null too often. At a fourfold spread of edge precisions, the false-positive rate inflates from five percent to over seven percent. Their fix is to whiten the data — rescale each edge by its precision — and that restores the test’s calibration.
Tom: And they’re honest about the limits there too. At an eightfold spread, even the whitened test drifts, and they recommend a permutation null instead. That’s the kind of careful, self-aware statistics I wish more papers had.
Jane: There’s also a really practical improvement in how they handle evidence gaps. You’d think missing data would make obstruction worse, but it’s the opposite. Deleting overlaps destroys cycles faster than it destroys the triangles that fill them, so the observed harmonic dimension goes down. A sparse evidence base looks more coherent than it actually is. That’s the dangerous direction — it means the prevalence of inconsistency in published evidence networks is understated for a structural reason, not just a statistical one.
Tom: So the improvements are: a mechanistic model instead of injected noise, a precision-whitened test, and a warning that gaps hide obstruction. What does that mean for someone actually deploying this?
Jane: That’s where Meng and Lu come in — I want to hear from them about what this looks like in practice.
Lu: Jane, I’ll jump in. The part that excites me most is the translation of a cohomology class into a data-collection instruction. This is the first time I’ve seen a topological obstruction used as a prescriptive guide for where to stratify next. That’s not just an abstention rule — it’s an experimental design tool.
Meng: And from an engineering standpoint, the pipeline is actually tractable. You build the nerve from your context overlaps, you populate the cochain from your measurements, you do two least-squares projections — that’s linear algebra, not deep learning — and you get a scalar harmonic energy. The compute cost is trivial. The hard part is eliciting the right context factorisation in the first place.
Jane: That’s a fair point. The paper admits the contexts are given, and choosing them well is the hard applied problem. But once you have them, the machinery is cheap and auditable.
Conclusion: Tom: So we’ve spent this episode on “Local verification cannot detect non-transportability,” and I want to pull it together before we hand off to the next paper.
Jane: The core message is that agentic AI systems — the ones that chain evidence across contexts — have a provable blind spot. Local verification, no matter how carefully implemented, cannot detect loop-level inconsistency. That’s not an implementation failure. It’s a structural property of the context space.
Tom: And the paper gives us the language to talk about it. The evidence decomposes into three parts: gradient, curl, and harmonic. The first two are fixable — recalibrate, re-measure. The third is not fixable at the current granularity. It’s an instruction to refine the covering, to stratify further, to collect data where the loop doesn’t close.
Jane: The simulations are controlled and honest. They include a correction of their own earlier claim, which I respect enormously. And the finance test case — where the null hypothesis is exactly known — gives the theory a clean validation you rarely get in biology.
Lu: If I can add one thing — the real-data validation the paper proposes is actually feasible with published data. Reconstructing the nerve from a published network meta-analysis and applying the precision-whitened F-test would place this framework directly against node-splitting, the established method, on its own ground. That’s a study that could be done this year.
Meng: And the engineering takeaway is that the compute cost is negligible. The bottleneck is domain expertise — knowing what the contexts are and where the overlaps live. That’s a human problem, not a GPU problem.
Tom: So where does this leave us? The paper doesn’t say local verification is useless. It says local verification is necessary but not sufficient. And the gap — the harmonic component — is measurable, it’s meaningful, and it points to a concrete action. That’s about as good as a negative result gets.
Jane: And it reframes abstention. Instead of a confidence score from a panel, you get a cycle witness — a specific loop where the evidence doesn’t close, with a recommendation for which stratification would fill it. For regulated industries, that’s a materially different artifact.
Tom: Alright, we’ve covered the theory, the method, the simulations, and the honest corrections. That’s “Local verification cannot detect non-transportability: Cohomological limits of context preservation in agentic reasoning.” Thanks for listening, and we’ll see you with the next paper.
Jane: Bye, everyone.
Suyash Mishra
cs.AI, cs.ET, cs.GT, cs.MA
Submitted: 2026-08-04
Comments: 17 pages, 10 figures. Includes an exact F-test for non-transportability with verified size and power, its precision-whitened generalisation, and a foreign-exchange case where the null hypothesis is known analytically rather than estimated
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 65/100
Key concepts
- Non-transportability / Harmonic Component
- A structural error in an AI agent's reasoning that only appears when evidence is chained around a large loop of contexts. Local checks miss this 'harmonic component,' which indicates the overall evidence network is inconsistent, even if every small segment looks fine.
- Local Verification
- The standard method for checking AI agents by verifying each step or context match individually (e.g., checking triples). The paper proves that this method is structurally blind to larger, loop-level inconsistencies in the evidence.
- Holonomy
- A term describing a shift or rotation that occurs when a claim is transported around a closed loop of contexts. It signifies an inconsistency—not noise—that results from the path taken through the evidence, rather than from local errors.
- Context Space
- The mathematical structure representing all possible settings or conditions (e.g., different diseases or cell types) where evidence is gathered. The paper analyzes how inconsistencies manifest within this space.
Terminology
Summary
Summary
This paper proves that local verification—the current state-of-the-art safeguard in agentic AI systems for checking context preservation—is structurally incomplete, and it introduces a cohomological method (Ks.etra) for detecting non-transportability of conclusions across contexts.
Core theoretical framework. The paper models context space as a site, evidence as a 1-cochain on the nerve of a covering, and agentic transport as path integration.
Specifically: Let X be a space of decision-relevant situations and U = Ui i∈I a finite covering by contexts.
An evidence cochain ω ∈ C1 assigns to each overlap ij the measured contrast in the quantity of interest between contexts i and j.
A claim assignment θ ∈ C0 is coherent with evidence if ω = δ0θ.
Main theorem (Theorem 5). θ̂t γ is independent of γ for every pair (a, t) if and only if ω ∈ im δ0.
For two paths γ, γ′ with common endpoints, the difference θ̂t γ − θ̂t γ′ = ⟨ω, γ − γ′⟩, which depends only on the class [ω] ∈ H1(N(U); R) whenever δ1ω = 0.
This establishes that inter-path disagreement is exactly holonomy.
Hodge decomposition (Theorem 7). C1 = im δ0 ⊕ H1 ⊕ im(δ1)T orthogonally.
Writing ω = g + h + c: "(a) g = δ0β is exactly the contamination produced by per-context instrument offsets β and is removed by recalibration against any single anchored context. (b) c ∈ im(δ1)T is the unique component with δ1c ≠ 0; it is precisely what a triple-overlap coherence check detects. (c) h satisfies δ1h = 0, so it passes every triple-overlap check, yet h ∉ im δ0, so by Theorem 5 it produces non-zero disagreement between some pair of reasoning paths."
Central negative result (Corollary 9). No family of simplex-supported consistency checks can distinguish ω from ω + h for h ∈ H1. Detecting H1 obstruction requires a statistic supported on a cycle basis of N(U), i.e. a genuinely global computation.
The paper clarifies: It is consistency checking specifically—the thing that makes verification cheap and specific—that cannot see the harmonic component.
Method (Ks.etra). The procedure: (1) site construction, (2) evidence assembly, (3) decomposition via least squares, (4) triage into recalibrate/re-measure/refine, (5) estimate by coboundary projection and abstain iff ∥h∥/√E exceeds a threshold calibrated to the decision tolerance.
Simulation results. Two domains: pharma RWE (contractible attribute space) and consumer credit (cyclic regime axis, carries dim H1 ≥ 1 by construction
). Key findings: Only harmonic energy predicts the error that survives optimal estimation; curl and gradient are statistically indistinguishable from zero
(Table 1, Claim B: harmonic ρ = 0.311 pharma, 0.471 credit; curl ρ = 0.032, −0.003). Gating on harmonic energy rather than total conflict lowers AURC by 0.029 and 0.039 (paired bootstrap, p < 0.001). The single-chain agent is worse than the context-blind pooled agent (MAE 1.82 vs 0.83). Chaining evidence compounds variance along the path.
Mechanism (Proposition 14). Non-transportability is therefore effect modification combined with overlap-specific population composition: Simpson's paradox on a network rather than on a single table.
At m = 0 (no effect modification), the harmonic energy is 1.52 × 10−15—machine zero.
Correction. "We previously reported that curl energy carries no information about irreducible error. That was an artefact of drawing the curl and harmonic components independently. In a mechanistic model a single latent cause—effect modification—generates both, so they co-vary, and curl is nearly as predictive as harmonic." Harmonic retains independent value (partial ρ = 0.19, p < 10−13). Two components may both predict error while demanding different responses, and it is the response that the decomposition is for.
Exact F-test. F = (∥h∥2/β1)/(∥c∥2/rank δ1) ∼ F(β1, rank δ1)
is an exact test for a global section under isotropic noise. Under unequal precision, the unweighted test is mildly anti-conservative—a 42% inflation of the false-positive rate at fourfold spread
(size 0.071 at nominal 0.05). The remedy is precision whitening: send ω ↦ Dω, δ0 ↦ Dδ0, δ1 ↦ δ1D−1.
Evidence gaps. dim H1 computed on an incomplete nerve is a lower bound on the obstruction. A sparse evidence base looks more coherent than it is
(Corollary 15).
Finance test case. In an arbitrage-free market, the log-quote cochain is exactly a coboundary.
The executability-aware nerve of a 15-currency complex reveals "five independent loops carrying P&L that no triangle check can detect. The harmonic detector
reaches AUROC 0.835 at 0.5 bp and 0.997 at 1 bp, whereas total-residual monitoring is still at 0.515 at 1 bp."
IFRS 9/SR 11-7 application. A PD model inventory over 20 contexts has 42 = 19 (calibration) + 6 (coherence) + 17 (transport)
degrees of freedom. Seventeen independent loops in the inventory can carry an incoherent PD story that no amount of pairwise or triple validation will detect.
Chaining one validated route degrades to 79.4 bp ECL misstatement at 0.30 log-odds incoherence, while projection stays flat at 5 bp.
Decision economics. Cohomological abstention is worth its complexity only where committing wrongly costs materially more than committing rightly gains—which is the regime of regulated therapeutic and credit decisions.
Limitations. "(1) There is no real data in this paper... (2) Correlated errors across overlaps are not handled... (3) Evidence gaps and obstruction interact, but not symmetrically—gaps hide obstruction... (4) We treat contexts as given... (5) Real-valued claims only."
Real-data validation proposed. "Reconstructing the nerve from a published network [meta-analysis], applying the precision-whitened F-test and comparing against the node-splitting results the original authors reported would place this framework directly against the established method on its own ground. Also:
construct the nerve over the 29 cell type × 5 disease contexts, populate ω from PINNACLE and TranscriptFormer contrasts, and test whether harmonic energy predicts (a) disagreement among MultiRoundDiscussion panellists and (b) which of the agent's confident answers are wrong."
Improvements for AI systems
Based on this paper, I can implement the following specific improvements to AI systems:
Implementation: After any multi-step reasoning chain, compute the evidence 1-cochain ω from all intermediate claims, decompose it via Hodge decomposition into gradient (g), curl (c), and harmonic (h) components, and abstain when ∥h∥ exceeds a calibrated threshold.
What the improved system can do: Detect when its conclusion is an artifact of the reasoning path taken, even when every local verification check passes. It will refuse to answer structurally unanswerable queries rather than confidently returning path-dependent results.
Abstract
Agentic AI systems routinely transport conclusions across biological, clinical and financial contexts, and the emerging safeguard is local verification: checking at each step that the entity is representable in the chosen tool, that parameters are compatible, and that outputs cohere with the plan. We prove this class of safeguard is structurally incomplete. Modelling a covering of context space by its nerve and evidence by a real-valued 1-cochain, an agent chaining evidence performs path integration: its conclusion is path-independent if and only if the cochain is exact, and disagreement between valid reasoning paths is exactly the holonomy of a first Cech cohomology class. Hodge decomposition partitions evidence conflict into a gradient part (calibration), a curl part (local inconsistency, visible at triple overlaps) and a harmonic part. Our central result is that no family of simplex-supported consistency checks can distinguish omega from omega+h for harmonic h, which nonetheless generates non-zero disagreement between valid paths; detection requires a statistic on a cycle basis. The resulting procedure, Ksetra, estimates by coboundary projection and gates abstention on the harmonic component, which we give a mechanism: it arises from effect modification combined with overlap-specific population composition, and vanishes to machine precision when effect modification is absent. The degrees of freedom of an evidence network partition into calibration, coherence and transport, yielding an exact F-test for the existence of a global claim; we quantify its distortion under unequal precision and supply the precision-whitened form that restores exactness. Foreign exchange, where the arbitrage-free null makes the cochain exactly a coboundary, serves as a calibration bench: the test is correctly sized, fires on loop arbitrage, and ignores triangular arbitrage.
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection