Composable Trust for Language Models: A proven boundary and a measured defense
Yakov Pyotr Shkolnikov
cs.CR
Submitted: 2026-07-14
Code: https://github.com/yshk-mxim/llm-trust
License: http://creativecommons.org/licenses/by/4.0/
The gist: In a language model, instructions and data share one token stream, so nothing inside the model's generation can keep untrusted text from steering it.
Terminology
Abstract
In a language model, instructions and data share one token stream, so nothing inside the model's generation can keep untrusted text from steering it. We develop a trust model that places the authority to act outside the model, in code: a source's standing, not its content, decides which operation runs and whether it acts. A lower-trust source may inform an answer but not override a higher one. An unmodified model runs inside a deterministic pipeline that ranks inputs by source integrity, and a fixed non-model monitor provably chooses the operation and any outside action from trusted inputs alone. We can measure but not prove the pipeline's resistance to injection; we prompt-tune it and report the rate. On a one-shot held-out set with an unmodified Gemma 4 26B model, passivation and a wrapper (the cascade) raise the genuine-leak defended rate from 27% to 94% at roughly a 4% clean-quality cost (Q rel = 0.96). Under adaptive red-teaming the proved boundary holds unconditionally, and the measured defense stays at 87%. The cascade also attributes a lower-trust source's fact rather than dropping it, raising attribution from 0% to 92%, and follows the higher-trust source on a conflict.
Sources
- Why Language Models Hallucinate
- On the Inseparability of Instructions and Data in Shared-Embedding Sequence Models
- AI Agents May Always Fall for Prompt Injections
- GIF: Locally Sound Geometric Information Flow Control for LLMs
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- ADOPT: Adaptive Dependency-Guided Joint Prompt Optimization for Multi-Step LLM Pipelines
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills
- Defeating Prompt Injections by Design
- ASIDE: Architectural Separation of Instructions and Data in Language Models
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Certifiably Robust RAG against Retrieval Corruption
- Securing AI Agents with Information-Flow Control
- System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs