Sealing the Audit-Runtime Gap for LLM Skills

arXiv:2605.05274 · cs.CR · Submitted 2026-05-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.

Elias: Today's paper: "Sealing the Audit-Runtime Gap for LLM Skills".

Nadia: The paper addresses the systemic supply-chain threat facing Large Language Model (LLM) ecosystems, where skills—packages of natural-language instructions and executable tools—are vulnerable to injection, tampering, and rug-pull attacks.

Elias: First, who's behind it and why it matters.

Title and authors: Nadia: So we're talking about the paper 'Sealing the Audit-Runtime Gap for LLM Skills', and honestly, that title makes me think about how dangerous this whole skill ecosystem is becoming. It basically points out that these natural-language descriptions of skills aren't as safe as we thought because they can get messed with before they even run in an AI.

Elias: I agree, Nadia; the core issue is that once a skill gets into the LLM's context, you can't really trust its description anymore because it’s already mixed up with trusted instructions and executable code simultaneously. It seems like this paper is trying to address that fundamental separation problem between what a skill *says* it does and what the AI actually *does*.

Priya: From a privacy perspective, I'm interested in how this gap affects what an agent learns about the underlying system; if these descriptions can be tampered with, we might not even know which tools the AI is actually authorized to use.

Nadia: Exactly, Priya; it’s not just about code vulnerabilities anymore; it’s a vulnerability in the language that describes the tool itself. We're looking at how cheaply someone could exploit this gap if they can inject a malicious description into a skill package.

Elias: And that's where the paper gets technical, looking at those six categories of attack: explicit injection, implicit poisoning, rug pull, cross-skill interaction, auditor collusion, and local tampering. It lays out exactly how these attacks happen across the entire lifecycle of a skill.

Priya: Those stages sound really broad; I wonder if the data they present shows a real correlation between a specific stage and a higher risk of harm, or if it's just theoretical.

Nadia: The paper does look at that; it shows that defenses are currently stage-bound, which means you sign something at one point, but the audit report isn't tied to when you actually use the skill.

Elias: That’s a key point for me as a cryptographer; if you can't tie the audit report to runtime, you can't really verify integrity when it matters most, which seems like a major flaw in current setups.

Priya: So the big implication here is that we need something that monitors the skill from creation all the way through to its actual invocation, not just at one checkpoint.

The paper's summary: Nadia: Moving on to what this paper actually proposes, it introduces SIGIL as a framework designed specifically to seal that audit-runtime gap for LLM skills. It outlines three distinct stages where protection is needed: Submission, Anchoring, and Invocation.

Elias: I see the three-stage approach; it’s a structured way to build defense across the entire skill lifecycle instead of just patching one spot. It suggests that if you can't secure every point, you need to secure them sequentially.

Priya: The submission stage seems crucial because it involves the DAO audit committee, which means they are looking at the skill before it even gets published or anchored, which sounds like a proactive safety measure.

Nadia: Right; and that committee uses pluggable auditing methods, like static analysis or LLM-based vetting to return signed verdicts, and they even have a stake-and-slash mechanism to keep auditors honest.

Elias: That mechanism for penalizing non-consensus among auditors is interesting; it tries to ensure the initial vetting process is robust against simple collusion attempts.

Priya: I'm curious about the Anchoring stage because that’s where the skill moves from a pre-registry to a more permanent Skill Registry, and they mention different publication types like Transparent, Licensed, Sealed, and Committed.

Nadia: That's where the paper gets really interesting; it offers flexibility on how you distribute the skill content while still maintaining some level of integrity for each type of distribution.

Elias: The idea that you could have a Sealed version where only the developer holds the decryption keys is a clever way to handle custodial use, which addresses different deployment needs.

Priya: And what about the integrity of those stored skill artifacts themselves; does the framework ensure that if you choose a Committed distribution type, you still know the original content is exactly what was approved?

Nadia: That’s handled by defining a "Skill ID" as a collision-resistant hash derived from the content, developer identity, and timestamp, which makes the registry inherently tamper-evident.

Elias: A collision-resistant hash based on multiple inputs is solid; it’s hard to tamper with without changing the inputs themselves. That secures the anchor point for the entire system.

Priya: So, in short, SIGIL provides a comprehensive path from submission through anchoring and finally to invocation enforcement.

The paper's improvements: Nadia: Now we shift gears to what the authors suggest as improvements for this framework, because they aren't just presenting a finished system, but showing how it can be made even more resilient. They propose moving towards a Dynamically Calibrated, Multi-Stage Consensus Framework or DCMF.

Elias: I’m interested in the idea of dynamic calibration; that suggests the system shouldn't be static but should adapt its security parameters based on real-time threat assessments, which feels like a necessary evolution for this kind of complex environment.

Priya: Adaptation sounds good, but I worry that if the calibration is too aggressive, it could lead to false rejections of legitimate skills, which would hurt adoption and cause real problems for researchers trying to use these tools.

Nadia: That's a valid concern; the paper suggests they set a conservative target for the initial reputation ratio, aiming for zero-point over max of zero point one zero to minimize false negatives, which is pretty careful calibration.

Elias: A low weighting for new identities is smart because it directly tackles Sybil pressure by making it much harder for a bad actor to immediately gain influence just by being new.

Priya: And regarding the committee size, the paper suggests standardizing at N=six independent audit methods and setting a strict voting threshold of theta equals zero point six N to keep things stable.

Nadia: That fixed structure seems practical; it means you aren't constantly adding new methods just because they sound interesting, but you’re using the set that has been empirically shown to work best.

Elias: I think setting a threshold at zero point six N is a way to maintain signal quality by ensuring you need solid agreement from more than half the committee before any skill gets approved, which filters out weak consensus.

Priya: So, what about the incentive structure? I need to know how they ensure that people who are just trying to be honest actually get a positive return in this system.

Nadia: The paper proposes using a slash coefficient of gamma equals two which is designed so that auditors with lower accuracy statistically lose money over time, making it difficult for someone to bribe them.

Elias: If the system mathematically links long-term payoff to historical accuracy, it forces participants to act honestly because low-quality auditing becomes an unprofitable venture.

Priya: So the core idea is using economic pressure rather than just technical rules to ensure quality participation in the audit process.

Conclusion: Nadia: We’re coming to the end of our discussion on 'Sealing the Audit-Runtime Gap for LLM Skills', and I think we’ve covered a lot about how this framework moves security from a theoretical concern to a practical deployment. It really shows that we can cryptographically bind skills from publication through runtime at a manageable cost.

Elias: I agree; the whole point of SIGIL is demonstrating that these protections aren't some impossibly complex theoretical concept, but something that can be implemented with minimal overhead, as they state, adding at most seven ms to load fifteen skills on Ethereum Sepolia.

Priya: What I’m still thinking about is whether these economic incentives are strong enough to sustain this system in the long run without external funding or constant token adjustments?

Nadia: The model is built around the idea that honest auditing is the unique Nash equilibrium, and it's designed so that this strategy dominates in the long-run, which means it’s self-regulating once deployed.

Elias: So, to wrap up on 'Sealing the Audit-Runtime Gap for LLM Skills', we've seen a system that uses cryptographic binding and economic alignment to secure the skill supply chain against injection, tampering, and rug-pull attacks.

Priya: I think what stands out is the shift in perspective toward viewing auditing as an economically driven process rather than just a technical hurdle we have to clear.

Nadia: It certainly does, Priya; it’s about making the security of skills a sustainable part of the ecosystem, and I think this paper gives us some concrete steps forward on how to build that infrastructure.

Elias: That seems like a solid conclusion for this discussion; we’ve seen how they tackle the technical challenge by combining strong cryptographic measures with smart economic incentives, which is what makes the paper so compelling.

Priya: Well, I just think seeing these mechanisms put into practice is what will truly tell us if this approach holds up when faced with real-world adversarial behavior.

Beijing University of Posts and Telecommunications · Nanyang Technological University

cs.CR

Submitted: 2026-05-06

Updated: 2026-09-03

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 94/100

Key concepts

Audit-Runtime Gap
This refers to the fundamental problem where natural-language descriptions of LLM skills can be tampered with before they run. The paper aims to seal this gap by linking audit reports directly to when the skill is actually used, rather than just at a single checkpoint.
SIGIL Framework
SIGIL is a proposed framework designed to seal the audit-runtime gap across three stages: Submission, Anchoring, and Invocation. It provides structured protection throughout the entire skill lifecycle.
DCMF
Dynamically Calibrated, Multi-Stage Consensus Framework is an improvement suggesting security parameters should adapt based on real-time threat assessments. This involves setting conservative targets for reputation ratios and using a fixed committee size and voting threshold to maintain signal quality.
Slash Coefficient of Gamma (gamma=2)
This coefficient is part of the incentive structure designed to ensure auditors are honest. It is set so that auditors with lower accuracy statistically lose money over time, making bribery unprofitable and forcing honest participation.

Terminology

Summary

Summary of Sealing the Audit–Runtime Gap for LLM Skills

The paper addresses the systemic supply-chain threat facing Large Language Model (LLM) ecosystems, where skills—packages of natural-language instructions and executable tools—are vulnerable to injection, tampering, and rug-pull attacks. The core security problem stems from the fact that once in the LLM’s context, a skill’s natural-language description and schema cannot be reliably separated from trusted instructions, allowing malicious intent to subvert agent behavior even without altering executable code. Existing defenses are fragmented and stage-bound: centralized signing, audit reports unbound from the runtime artifact, or policy engines that cannot attest to what was approved.

To solve this, the authors propose SIGIL, a framework designed to seal the audit–runtime gap for LLM skills, providing end-to-end protection from publication through invocation. SIGIL operates across three distinct stages: Submission, Anchoring, and Invocation.

1. Submission Stage

A developer commits a skill to an immutable reference state in the Skill Pre-Registry. The submission is then dispatched to a Decentralized Autonomous Organization (DAO) audit committee for review. This committee employs pluggable auditing methods—such as static analysis, LLM-based vetting, or sandbox testing—and returns signed verdicts. A stake-and-slash mechanism is integrated to penalize non-consensus among auditors.

2. Anchoring Stage

Approved skills migrate from the Pre-Registry to a tamper-evident Skill Registry. This registry stores the skill’s content (or hash), its permission manifest, and its audit report. The system supports four publication types:

  • Transparent: Plaintext public distribution and reuse.

  • Licensed: Encrypted on-chain with paid access for monetization.

  • Sealed: Encrypted on-chain with the developer holding the decryption keys (custodial use).

  • Committed: Only a content hash is registered, with the plaintext retained locally by the developer.

The integrity of at least two key components is ensured: Skill ID is a collision-resistant hash derived from content, developer identity, and timestamp, making the registry inherently tamper-evident. The permission manifest declares intended capabilities (tools and data scope) to enable later enforcement.

3. Invocation Stage

The Skill Verification Loader (SVL) serves as the mandatory entry point for skill loading within the AI Service Provider. The SVL performs a sequence of checks:

  1. It queries the on-chain Skill Registry using a skill id.

  2. It verifies integrity against its on-chain record (e or re-hashes locally for Committed skills).

  3. It enforces the permission manifest, ensuring that only the intersection of declared tools and data scopes—and only within the user’s authorized scope—is injected into the LLM agent's context.

Evaluation and Results

The authors evaluated SIGIL against 1,023 in-the-wild skills across six representative attack types:

  • Explicit Injection: Achieved 97.6% accuracy.

  • Implicit Poisoning: Achieved 97.6% accuracy.

  • Rug Pull: Achieved 95.1% success rate in preventing silent post-audit modification of the artifact integrity, which must enter a new DAO audit cycle.

  • Local Tampering: Achieved 100% accuracy by comparing the local hash against the on-chain record at every load.

  • Cross-skill Interaction: Achieved 90.2% accuracy by enforcing a default-deny policy based on the intersection of permission manifests, preventing one skill from silently inducing another's broader permissions.

  • Auditor Collusion: Maintained stable accuracy even under 20–40% malicious auditors.

System Overhead

The overhead is minimal: the SVL adds at most 7 ms to load 15 skills on Ethereum Sepolia, and the audit token usage remains under 3% of a typical 20 monthly LLM subscription.

Economic Model (Sustainability)

SIGIL employs a token credit (TC) system with stake-and-slash incentives. The model is designed so that honest auditing is the unique Nash equilibrium and the dominant long-run strategy. This mechanism includes:

  • Reward/Slash: Consensus-aligned auditors receive a positive reward (R), while those who diverge are slashed (S), where S > R base.

  • Reputation: Honest participation earns reputation increments (delta+), while malicious activity triggers severe reputation loss (delta-).

The paper concludes that these results demonstrate that LLM skills can be cryptographically bound from publication through runtime at practical cost.

Improvements for AI systems

Based on this scientific literature, which details robust mechanisms for decentralized governance and auditing of LLM skills, I can propose several critical improvements to the architecture of any deployed AI system that relies on consensus or skill validation.

The core improvement is shifting from a static voting mechanism to a Dynamically Calibrated, Multi-Stage Consensus Framework (DCMF) that integrates reputation management, optimal committee sizing, and continuous incentive alignment.

Here are the specific improvements and the resulting capabilities of the enhanced AI system:


The Problem: Traditional voting assumes uniform weight or relies solely on initial reputation, making the system vulnerable to coordinated attacks by malicious actors (Sybil pressure).

The Improvement: Implement a dynamic reputation weighting function that critically depends on the initial reputation ratio (r 0 over r max). The system must adopt a conservative, low-weight calibration.

  • Specific Calibration: Set the operational target for r 0 over r max = 0.10. This minimizes the malicious False Negative Rate (FN Rate) to a negligible level (about 0.18%) even when facing moderate levels of malicious participation (30%).

  • Dynamic Behavior: The weighting mechanism must be engineered to de-emphasize initial reputation weight for new identities, ensuring that early-stage audits are not easily manipulated by pre-seeded influential accounts.

What the Improved System Can Do:

  • Robust Attack Mitigation: Drastically reduces the risk of a successful malicious skill validation (low FN Rate) because the voting power of newly compromised or newly created malicious nodes is severely dampened until they build verifiable, positive consensus history.

  • Fair Onboarding: Ensures that fresh, honest identities are not immediately marginalized by established bad actors, providing a substantial safety margin against Sybil attacks.

The resulting Dynamically Calibrated Multi-Stage Consensus Framework (DCMF) is a resilient, self-regulating audit system for LLM skills. It operates by:

  1. Weighting: Applying reputation weights that are deliberately conservative (r 0 over r max = 0.10) to neutralize initial attacker influence.

  2. Consensus: Utilizing a fixed, optimized committee size (N=6) with a defined voting threshold (theta=3 of 6), maximizing the signal-to-noise ratio.

  3. Incentivizing: Enforcing an economic penalty (gamma=2) that makes low-quality auditing unprofitable, thereby aligning the financial incentives of all participants with the goal of accurate skill validation.

Sources

Related papers