Approval Integrity and Recovery in LLM Answer Publication
cs.CR
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 21 pages, 15 tables. Reproducibility code, losslessly compressed experimental results, and integrity-verifying restoration scripts included as ancillary files
Code: https://github.com/beir-cellar/beir
License: http://creativecommons.org/licenses/by/4.0/
The gist: Publication integrity in LLM systems requires binding approved content to its current authorization context.
Terminology
Abstract
Publication integrity in LLM systems requires binding approved content to its current authorization context. We examine exact-content binding, authorization freshness and checkpoint recovery in Lightcap's publication enforcement mechanism. On 900 independently human-annotated RAGTruth responses from 150 source tasks, three dated Ministral models and a same-model direct-grounding baseline yield 3,600 assessments. The production response-act checker instantiated with 14B accepts 291 of 302 unsupported-labelled answers; the direct baseline accepts 41. Supported-answer retention is 95.2% and 66.9%, respectively. An exact promotion-correction identity tracks error through 100 chronological 3B-14B-8B-14B answer trajectories. Among 65 initially approved answers, the final stateful recheck-recovery policy increases exact-match error by 9.23 percentage points relative to the initial checkpoint (95% article-clustered interval [-1.72, 19.61]). Controlled evidence-fingerprint changes expose asymmetric freshness enforcement between publication and recovery. A separate BIPIA prompt-injection experiment records zero target insertions among 266 valid editor outputs. External Hugging Face calibration experiments transfer retrieval models from ArguAna to SciFact and NFCorpus, and diagnostic decision rules from Thunderbird to BGL, distinguishing probability calibration from ranking changes. The measurements separate semantic false approval, stale authorization and recovery-induced error at executable publication boundaries.
Sources
- Indirect Prompt Injections: Are Firewalls All You Need, or Stronger Benchmarks?
- Securing AI Agents with Information-Flow Control
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- Defeating Prompt Injections by Design
- SPA: Securing Persistent LLM Agents Across Queries with Plan-First Information-Flow Control
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Efficient Code Embeddings from Code Generation Models
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Mind the Gap: Time-of-Check to Time-of-Use Vulnerabilities in LLM-Enabled Agents
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- SQuAD: 100,000+ Questions for Machine Comprehension of Text
- Toolformer: Language Models Can Teach Themselves to Use Tools
- BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
- AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents
- CAVA: Canonical Action Verification and Attestation for Runtime Governance of Agentic AI Systems
- Proof-Carrying Agent Actions: Model-Agnostic Runtime Governance for Heterogeneous Agent Systems
- ReAct: Synergizing Reasoning and Acting in Language Models
- Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
- Loghub: A Large Collection of System Log Datasets for AI-driven Log Analytics
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs