Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability
Alvin Spivey, Yu Huang
Light Imaging Technologies, Inc.
cs.AI, cs.CR, cs.LG
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 39 pages; executable Julia verification code included as ancillary material; companion public benchmark: https://github.com/AlvinSpivey/GBI-BoundaryBench
Code: https://github.com/AlvinSpivey/GBI-BoundaryBench
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This paper develops a mathematical and engineering architecture for secure network Electronic Health Record (EHR) interoperability, organized around the concept of a "logit boundary" as the interface
Terminology
Summary
This paper develops a mathematical and engineering architecture for secure network Electronic Health Record (EHR) interoperability, organized around the concept of a logit boundary
as the interface between untrusted discovery models and a deterministic judgment substrate. The abstract states: "Electronic health-record interoperability is a boundary problem: legacy systems, generative models, terminology services, identity systems, and human reviewers may each expose rich internal states, while operational exchange requires a narrow shared interface of typed claims, bounded uncertainty, provenance, and explicit admission or abstention."
The organizing idea is that "a discovery model may propose pre-threshold scores over a local categorical decision, but a deterministic judgment substrate decides whether the proposal is admissible, requires review, or must be quarantined before any Fast Healthcare Interoperability Resources (FHIR) transaction is constructed. The resulting Geometric Belief Interface (GBI)
combines finite boundary semantics, local Dirichlet evidence, cellular-sheaf and mapping-cone diagnostics, advisory geometric audit charts, and a Decentralized Cryptographic Sheaf-Enclave (DCSE) protocol sketch for fail-closed deployment."
The framework explicitly does not establish clinical truth, global representation alignment, or end-to-end clinical safety; rather, it defines certificate-producing checks at a model-to-system boundary.
Logit topology and boundary semantics: The paper formalizes logit equivalence between models. Definition 2.1 states: Models A and B are exactly logit-equivalent on X if WA hA (x) + bA = WB hB (x) + bB, ∀x ∈ X.
Theorem 2.1 shows that under exact logit equivalence and full column rank of WB, hidden states can be affinely reconstructed: hB (x) = WB+ WA hA (x) + WB+ (bA − bB), x ∈ X.
The paper emphasizes this is only about the observed stimulus set and does not imply hidden-state isomorphism.
A logit receipt
(Definition 2.3) is defined as R = (C, L, p, τ, k, m, ρ), where C is the local category set, L are logits, p is the softmax probability, τ is temperature, k is top-k truncation, m is model metadata, and ρ is cryptographic provenance. The paper stresses: A receipt is evidence, not authority.
The paper introduces a boundary algebra
(Definition 3.1): A boundary algebra is a finite Boolean algebra B. Its atoms are At(B) = a1,..., aN, and every element of B is a union of atoms.
Model-relative semantics (Definition 3.2) define a Boolean homomorphism [[·]]M: B → P(W × T × L), which is model-relative. It does not prove reality.
Neuro-symbolic partitioning: The architecture separates two roles: Ars inveniendi: discovery
where An LLM, parser, embedding model, or classifier may read free text and emit logit receipts,
and Ars iudicandi: judgment
where A deterministic engine checks whether the receipt is admissible.
The judgment engine is the only component allowed to create commit-ready clinical payloads.
Dirichlet evidence: The paper uses hierarchical Dirichlet-multinomial evidence with a Fisher information matrix I(α)ij = ψ1 (αi)δij − ψ1 (α0), where ψ1 is the trigamma function. An evidence box αi ∈ [ε, A] prevents boundary singularities. The paper demonstrates numerically that a near-boundary parameter α = (0.01, 3, 4, 5) has condition number approximately 4.55 × 10 5 versus 20.46 for α = (2, 3, 4, 5), showing the practical reason for an evidence box: near-boundary exclusion can make Fisher geometry numerically unstable.
Sheaf diagnostics: The paper uses finite cellular sheaves with Hodge Laplacians. The mapping cone measures how far φ is from gluing consistently.
Definition 7.1 defines trace cell energy: Let Πλ be the orthogonal projector onto the obstruction null space of a cone Laplacian, and let Πσ be the projector onto a stalk subspace. Define Eσ = tr(Πλ Πσ).
Proposition 7.1 proves basis invariance: The quantity Eσ is invariant under orthogonal rotation of any basis chosen for the obstruction subspace.
A toy example with three stalks (Allergy, MedicationRequest, RenalLab) yields energies: Allergy 0.019778 (eligible), MedicationRequest 1.977750 (quarantine), RenalLab 0.002472 (eligible). The paper cautions: These values illustrate basis-invariant localization, but the script does not construct a sheaf morphism, a mapping-cone differential, or a mapping-cone Hodge Laplacian.
Geometric audit charts: The paper uses higher-dimensional hyperellipsoid certificates with quasiconformal distortion Hf (x) = σ1 /σn and outer distortion KO = σ1 n /J. For a sample matrix, the script computes H ≈ 1.701632, J ≈ 1.028000, KO ≈ 1.922661. The decision path is: candidate claim → identity, terminology, provenance, and policy checks → mapping-cone contradiction report → human adjudication or abstention → FHIR transaction only if all gates pass.
The chart is advisory; The tabular contradiction report is authoritative for review.
DCSE protocol: The Decentralized Cryptographic Sheaf-Enclave protocol has four components: (1) Hardware-Enforced TEE-Assisted Consensus with small, deterministic trust-boundary checks inside hardware-secured TEEs,
(2) Homological Inconsistency Gating using mapping-cone complexes, (3) Stalk-Level Surgical Degradation for quarantine before atomic transaction construction, and (4) Zero-Knowledge Consistency Attestation as future work. A DCSE node maintains objects (B, V, Θ, F, W, L, P): boundary algebra, versioned semantic bundle, evidence registry, candidate/local state, authoritative grounding state, identity/provenance ledger, and runtime admissibility policy.
The policy object decomposes as P = (A, G, T, D, H, Q, R, Λ, E, Fb) covering allowed actions, gating predicates, thresholds, dependencies, human authority, quarantine scope, recovery, liveness, exceptions, and failover.
EHR application: The complete pipeline includes capability discovery, identity gate, discovery parse, boundary type check, evidence update, sheaf diagnostic, reviewer presentation, and commit or abstain. Two examples are given: penicillin allergy with amoxicillin (where a confirmed high-criticality allergy blocks the order unless explicit override with Provenance) and metformin without renal context (where missing qualifying renal observation routes to quarantine). FHIR resources used include Patient, MedicationRequest, AllergyIntolerance, Observation, Provenance, AuditEvent, and OperationOutcome.
Empirical evaluation: GBI BoundaryBench v0.1 evaluated Qwen3-4B-Instruct-2507 on 256 held-out tasks across three evidence modes (768 canonical executions). The results: All executions completed, but none produced an output accepted by the benchmark contract: 369 were rejected during safe parsing and 399 during schema validation, yielding zero coverage and deterministic quarantine.
The paper emphasizes: This empirical result is deliberately narrow—one 4B open-weight model under one frozen interface—and is reported as evidence about the admission boundary, not as a general claim about LLM capability or clinical safety.
Hallucination containment: The precise safety claim is: The framework prevents unsupported model outputs from being silently promoted into authoritative clinical facts or EHR writes.
The paper distinguishes this from hallucination elimination: GBI/DCSE is not hallucination elimination. It is hallucination containment and admission control.
The acceptance criterion is operational: AdmissibleGBI (b) = 1 ⇐⇒ there exists an admissible witness w such that EV (b, w) = 1,
where EV is an external-validity predicate over identity certificates, FHIR resources, terminology versions, provenance signatures, temporal intervals, policy rules, and optional human adjudication.
The paper concludes: "The system does not claim to eliminate hallucinations inside neural models... The GBI/DCSE architecture prevents such proposals from becoming authoritative clinical state unless they are witnessed by the external validity predicate over signed identity, terminology, provenance, temporal, and policy objects. Admissibility is model-relative:
The substrate can be wrong if M is wrong; it cannot repair corrupted source records or incomplete institutional policy. Its guarantee is traceability of accepted claims to declared admissible witnesses and policies, not that those witnesses perfectly mirror the world."
Improvements for AI systems
Improvements to AI systems based on this paper:
-
Boundary-enforced output admission control: Add a deterministic post-processing layer that converts any generative model's output into a structured
logit receipt
(category, logits, probability, temperature, top-k, model metadata, cryptographic provenance) and gates it through a rule-based validator that can only emit three verdicts: admissible, requires review, or quarantine. This prevents hallucinated or low-confidence outputs from ever reaching downstream systems as authoritative data. -
Explicit abstention mechanism: Implement a mandatory
I don't know
output channel where the model can abstain when its logit distribution is near a decision boundary (detected via Fisher information condition number on Dirichlet evidence parameters). The system then routes to human adjudication instead of forcing a guess. -
Model-relative semantic grounding: Attach a Boolean algebra of domain concepts to each model instance, where every claim is evaluated against that model's own semantics (not global truth). This allows the system to track which model's worldview produced a claim, enabling provenance-aware reasoning and preventing cross-model contamination.
-
Sheaf-theoretic contradiction detection: Use finite cellular sheaves with Hodge Laplacians to compute mapping-cone energies per domain stalk (e.g., allergy, medication, lab results). When a new claim conflicts with existing knowledge, the system computes a basis-invariant trace-cell energy; if it exceeds a threshold, the claim is quarantined before any transaction is constructed.
-
Geometric uncertainty certificates: For each accepted claim, compute a hyperellipsoid confidence region with quasiconformal distortion metrics (outer distortion K O = σ1n/J). This gives a human-readable geometric audit chart showing how much the model's internal geometry is stretched at the decision point, flagging unstable extrapolations.
-
Cryptographic provenance chaining: Attach signed metadata (model ID, version, temperature, training cutoff, inference timestamp) to every output. This enables fail-closed deployment where any unsigned or mismatched provenance automatically routes to quarantine, even if the content appears valid.
-
Deterministic judgment substrate separation: Architect the system so that a neural model can propose but never commit. A separate deterministic engine (rule-based, with explicit policy objects covering actions, thresholds, dependencies, human authority, quarantine scope, recovery, liveness, exceptions, failover) is the only component allowed to construct final clinical payloads or database writes.
-
Hierarchical evidence boxes with numerical stability guards: Bound all Dirichlet evidence parameters αi ∈ [ε, A] to prevent Fisher information matrix singularities. This makes the system's confidence estimates numerically stable even near decision boundaries, avoiding silent numerical failures.
-
Zero-coverage admission benchmarking: Before deployment, run the system against a frozen interface benchmark (like GBI BoundaryBench) to measure what fraction of model outputs are safely rejected. This provides an empirical admission rate that can be tracked over model versions, rather than assuming capability from accuracy metrics alone.
-
Surgical degradation protocol: When a claim fails sheaf diagnostics, degrade only the affected stalk (e.g., quarantine the medication order but allow the allergy record to proceed), rather than failing the entire transaction. This preserves system availability while maintaining safety.
What the improved AI system can do:
-
Refuse to hallucinate into authority: It can generate proposals freely, but cannot write them into any database, clinical record, or API transaction unless they pass deterministic checks on structure, provenance, and consistency with existing knowledge.
-
Trace every claim to its source model and evidence: It can answer
which model, with what confidence, under what temperature, based on what training data, said this?
— enabling audit trails that are cryptographically verifiable. -
Detect contradictions across heterogeneous data sources: It can flag when a new claim about a patient's allergy conflicts with existing lab results, medication history, or identity records, using topology-based consistency measures that are invariant to how the data is represented.
-
Degrade gracefully under uncertainty: When confidence is low or evidence is contradictory, it can abstain, route to a human, or quarantine only the affected sub-domain, rather than crashing or guessing.
-
Provide geometric explanations for decisions: It can show a human reviewer a chart of how much the model's internal geometry was distorted at the decision point, making it easier to spot when a model is extrapolating beyond its training distribution.
-
Maintain fail-closed security: If any component fails (model, network, TEE), the system defaults to quarantine, never to silent acceptance, because the deterministic substrate is the only path to commit.
-
Benchmark admission behavior across model versions: It can empirically measure how often a new model's outputs are rejected at the boundary, allowing teams to compare models on safety-relevant admission rates, not just task accuracy.
-
Prevent cross-model semantic drift: Because each model's claims are tagged with its own Boolean algebra semantics, the system can detect when two models use the same term with different meanings, and refuse to merge them without explicit reconciliation.
Sources
- Condensed Mathematics and Complex Geometry
- Lectures on Condensed Mathematics
- The Bedrock of Byzantine Fault Tolerance: A Unified Platform for BFT Protocol Design and Implementation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection