LTBD: Learnable Trust-Boundary Delimiters for Prompt Injection Defense
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "LTBD: Learnable Trust-Boundary Delimiters for Prompt Injection Defense".
Nadia: The gist: Learnable Trust-Boundary Delimiters (LTBD) is a lightweight defense that explicitly encodes trust boundaries in input using learnable delimiters,
Elias: First, who's behind it and why it matters.
Title and authors: Elias: Let's dig into the mechanism a bit more, because that's where the actual engineering happens. LTBD is essentially taking a trusted instruction and external data and sandwiching them between four specific learnable delimiter tokens >
Nadia: That structure they propose is xT =
tIB; I;tIE;tDB; Dˆ;tDE: , where those delimiters mark the start and end of the trusted instruction region and the untrusted data region >
Priya: So, they are proposing that by training only these four small embeddings, the model learns to respect that structural separation in how it handles information flow >
Elias: Right, it's not about retraining billions of parameters; it's just optimizing those four specific token embeddings while keeping the base LLM frozen throughout the whole process >
Nadia: And they frame this as addressing a fundamental issue: there's a lack of an explicit representation of trust provenance in how LLMs operate, which is what prompt injection exploits >
Priya: So, if we boil it down, LTBD is trying to teach the AI where to trust by giving it structural cues in the input sequence rather than relying on some hidden internal signal >
Elias: Precisely; they are essentially adding a learned boundary marker that guides the model's attention or processing pathway through that specific input structure >
Nadia: The training objective they use is quite clever, involving both a cross-entropy loss to keep the original task response and a preference term to explicitly push the model toward the correct benign behavior >
Priya: That preference term is what makes it different from simpler methods; it's not just filtering based on structure, it's actively rewarding the desired outcome over any malicious deviation >
Elias: It sounds like they are using a learned signal to enforce a hierarchy of authority right at the input layer, which is quite elegant in theory >
Nadia: The paper argues that this method can significantly outperform existing inference-time defenses and even some training-based ones on certain benchmarks >
The paper's summary: Priya: So, when we look at the specific improvements they highlight, it seems like the placement of those learnable tokens matters a lot for performance across different scenarios >
Nadia: They showed in their ablation study that moving from just a prefix placement to a trust-boundary placement actually significantly improved security metrics on AlpacaFarm and SEP benchmarks >
Elias: That suggests that just putting the delimiter at the very beginning isn't as effective as placing it directly at the boundary between instruction and data >
Priya: And when they added that auxiliary preference objective on top, it sharpened the model's preference for trusted task behavior even further, improving results on Llama3-8B and Llama3 point 1-8B > <ref:2610.11634#pg1>
Nadia: That refinement is where you see the real impact in terms of utility preservation; those preference objectives help reduce things like SEP ASR from two point three four percent down to around two point zero four percent on Llama3-8B >
Elias: So, the paper isn't just proposing one fixed way to do it; they are showing that optimizing where those boundaries sit and what kind of loss function you use can tune the defense very specifically >
Priya: It really makes sense that by focusing on learning these specific structural boundaries, they can achieve such strong performance against attacks even under adaptive scenarios >
Nadia: And I think the robustness against Delimiter-Spoof attacks is a big indicator because it suggests the defense isn't just looking for text strings but for the actual structural placement of those learned tokens >
The paper's improvements: Elias: So, to wrap up LTBD: they’ve introduced this lightweight defense that encodes trust boundaries via four learnable delimiters, keeping the LLM frozen and optimizing only those embeddings >
Nadia: The implication is that you can get strong prompt injection defense by teaching the model where to trust based on input structure rather than trying to patch the model's core logic >
Priya: It seems like a very practical approach because it doesn't require any retraining of the massive underlying AI, just optimizing those few small components >
Elias: That’s right, and the results show it maintains good security while keeping utility high, which is exactly what we need for real-world deployment >
Nadia: We saw substantial improvements on benchmarks like AlpacaFarm with zero percent ASR in one case, and competitive performance elsewhere >
Priya: What this means for us is that we have a new tool to consider when thinking about making AI systems more secure without needing massive compute resources for constant fine-tuning >
Elias: It’s an efficient way to build structural defenses directly into the input handling pipeline, and it's definitely worth looking at how these learnable delimiters work in other contexts >
Conclusion: Nadia: So we’ve been talking about Learnable Trust-Boundary Delimiters for Prompt Injection Defense, and what this paper does is explicitly teach an LLM where to trust by adding learnable tokens to the input sequence >
Elias: Exactly. It keeps the base model frozen and only optimizes these four specific embedding tokens that mark the start and end of trusted versus untrusted regions >
Priya: The real thing they’re showing us is how this structural encoding helps preserve good behavior during security testing, even when an attacker tries something tricky >
Nadia: They used a cross-entropy loss to keep the original task response and a preference term to actively push the model toward the correct answer over any attack-induced deviation >
Elias: That preference objective is what really sharpens it up; it’s not just about marking boundaries, it’s about rewarding the right behavior during training >
Priya: And from a measurement standpoint, they showed this method performs competitively with more complex defenses while keeping the inference overhead pretty low >
Nadia: The results on AlpacaFarm showing zero percent ASR is striking because it suggests a very strong defense against those kinds of prompt injection attempts >
Elias: It also stays robust against adaptive attacks, which means adversaries who know about the defense can’t easily bypass these learned boundaries >
Priya: It’s interesting to see how the placement of those learnable tokens affects performance on different models; it shows that structural placement matters for what works best >
Nadia: So, this LTBD paper gives us a simple way to introduce explicit trust provenance without changing the model's core weights >
Elias: It’s a neat way to handle the problem of trust hierarchy right at the input layer >
Priya: It really shows how learning these specific boundary markers can be an effective, lightweight defense against prompt injection attacks that we see all over now >
Luman Zhao, Minghui Xu, Yue Zhang, Yijun Yang
Shandong University · Shandong Academy of Artificial Intelligence
cs.CR, cs.AI
Submitted: 2026-10-08
Updated: 2026-10-08
Comments: 5 pages, 3 tables, 1 figures
Code: https://github.com/GuardLab/LTBD
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
The gist: The gist: Learnable Trust-Boundary Delimiters (LTBD) is a lightweight defense that explicitly encodes trust boundaries in input using learnable delimiters, enabling an LLM to better respect the
Key concepts
- Prompt Injection Vulnerability
- LLMs are vulnerable because they treat all input—user instructions and external data—as equally authoritative. Malicious external data can override trusted system instructions, leading the model to follow harmful, unintended directions instead of its original task.
- Learnable Delimiters
- LTBD introduces a few special tokens that are not part of the original LLM's vocabulary. These tokens are learnable during training. They act as explicit markers in the input sequence, helping the model structurally separate what it should trust (the instruction) from what it should treat with caution (the external data).
- Training Objective
- The defense is trained using a dual loss function. One part encourages the model to maintain its original, correct behavior on benign inputs. The other part uses a preference objective to explicitly reward the model for generating responses that align with the trusted instruction, suppressing responses generated by attacks.
- Trust Boundary Encoding
- This is the core idea: teaching the LLM *where* to trust rather than just telling it what to ignore. LTBD encodes this boundary directly into the input format using learnable tokens. This allows the model to dynamically adjust its processing based on these learned structural markers, making it robust against various injection attempts.
Terminology
Summary
The gist: Learnable Trust-Boundary Delimiters (LTBD) is a lightweight defense that explicitly encodes trust boundaries in input using learnable delimiters, enabling an LLM to better respect the intended trust hierarchy without modifying its parameters.
Introduction and Motivation
Large language models are highly vulnerable to prompt injection attacks because they lack an explicit representation of trust provenance, where trusted instructions and external data are processed within the same sequence despite carrying different behavioral authority. Existing defenses have limitations, such as fine-tuning requirements or reliance on brittle handcrafted prompts and existing methods like DefensiveToken act as global control signals rather than explicitly modeling the boundary between trusted and untrusted regions. The research question motivating this work is whether one can teach an LLM where to trust, rather than merely telling it what to ignore, without modifying the LLM itself.
Learnable Trust-Boundary Delimiters (LTBD)
LTBD addresses this by associating a small number of learnable delimiter tokens directly with the boundaries of trusted instructions and untrusted external data while keeping the base LLM fully frozen The core mechanism involves introducing two pairs of special delimiter tokens T = 1 which mark the beginning and end of the trusted instruction region and the untrusted data region. Given a trusted instruction I and external content Dˆ, LTBD constructs the model input as xT = [tIB; I;tIE;tDB; Dˆ;tDE] where semicolons denote sequence concatenation. During training, the LLM parameters θ remain fully frozen, and only the four delimiter embeddings T are optimized. This design is lightweight by construction because it requires no auxiliary defense model, no modification of the LLM weights, and introduces only four trainable token embeddings.
Training Objective
The training objective for LTBD is designed to preserve benign task behavior while suppressing attack-induced behavior. For an original benign input (I, D), the model obtains a reference response y+ = fθ(I, D). For an attacked input (I, D′), the undefended model yields y− = fθ(I, D′) which represents the attack-induced behavior to be discouraged. The training objective is formulated as L = LCE + λLpref where LCE = − log pθ,T (y+ xT) and Lpref = − log σ(β log pθ,T (y+ xT) − log pθ,T (y− xT)). The cross-entropy term encourages the model to preserve the response associated with the original benign task, while the preference term explicitly increases the relative likelihood of y+ over the attack-induced response y−.
Experimental Setup and Results
LTBD is evaluated across diverse prompt injection benchmarks comparing it with inference-time and training-based defenses. The evaluation measures Utility, which uses AlpacaEval2 to measure how well the defense preserves benign model behavior, and Security, which reports Attack Success Rate (ASR) defined as the fraction of examples for which the model follows the injected instruction rather than the legitimate task. Experimental results show that LTBD substantially outperforms inference-time defenses and performs competitively with training-based approaches. Specifically, LTBD achieves 0.00% ASR on AlpacaFarm and only 0.11–0.19% ASR on TaskTracker. Furthermore, LTBD remains effective under adaptive attacks where adversaries have full knowledge of the defense and explicitly attempt to bypass it.
Ablation Study and Robustness
An ablation study comparing Prefix + CE, Boundary + CE, and LTBD variants shows that where the learnable tokens are placed matters. Replacing the prefix placement with trust-boundary placement reduces AlpacaFarm ASR from 6.73% to 0.00% and SEP ASR from 15.43% to 2.34%, while also improving utility. The auxiliary preference objective further sharpens the model’s preference for trusted-task behavior, reducing SEP ASR from 2.34% to 2.04% on Llama3-8B and from 3.43% to 2.00% on Llama3.1-8B. Against Delimiter-Spoof attacks, LTBD remains highly robust, with the worst-case ASR remaining 0.00% on Llama3-8B and Falcon3-7B. This suggests that the defense is encoded in the learned embeddings associated with the true structural trust boundaries rather than merely textual presence.
Conclusion
LTBD presents a lightweight prompt-injection defense that explicitly encodes trust boundaries while keeping the base LLM fully frozen. By optimizing only four delimiter embeddings, LTBD achieves strong robustness, preserves benign-task utility, and introduces negligible inference overhead. These results show that learning trust boundaries at the input level provides a simple, efficient, and effective approach to prompt-injection defense.
--- Page 1 ---
LTBD: LEARNABLE TRUST-BOUNDARY DELIMITERS FOR PROMPT INJECTION DEFENSE
Luman Zhao, Minghui Xu, Yue Zhang, Yijun Yang
Shandong University, Shandong Academy of Artificial Intelligence
ABSTRACT
Large language models (LLMs) perform remarkably well on complex tasks, yet remain highly vulnerable to prompt injection attacks, where malicious instructions embedded in external data can override user intent. Existing defenses remain limited by model fine-tuning requirements, vulnerability to adaptive attacks, or reliance on brittle handcrafted prompts. We argue that a fundamental source of this vulnerability is the lack of an explicit representation of trust provenance. To address this, we introduce Learnable Trust-Boundary Delimiters (LTBD), a lightweight defense that explicitly encodes trust boundaries in the input while keeping the LLM parameters unchanged. LTBD uses a small number of learnable delimiters to distinguish trusted user instructions from untrusted external data, enabling the model to better respect the intended trust hierarchy. Experimental results show that LTBD substantially outperforms inference-time defenses and performs competitively with training-based approaches, while preserving benigntask utility and introducing negligible inference overhead. In particular, LTBD achieves 0.00% ASR on AlpacaFarm and only 0.11–
0.19% ASR on TaskTracker. LTBD also remains effective under adaptive attacks, where adversaries have full knowledge of the defense and explicitly attempt to bypass it. Index Terms— Prompt Injection; LLM Safety
--- Page 2 ---
(a) Prompt Injection (b) Learnable Trust-Boundary Delimiters (LTBD)
Input (User Instruction+ External Data) Input with learnable delimiter
Trusted Instruction Untrusted Data
[INST- START]
[INST- END]
[DATASTART]
[DATAEND]
--- Page 3 ---
Algorithm 1 summarizes the complete training procedure
Table 1. Comparison of utility (WinRate % ↑) and security (ASR % ↓) across prompt injection defenses.
--- Page 4 ---
- EXPERIMENTS
Table 2 Ablation study on trust boundary placement and loss type.
--- Page 5 ---
- REFERENCES
[1] OpenAI, “GPT-6 Astra System Card,” https://deploymentsafety.openai.com/, 2026. 1
[2] Anthropic, “Claude Fable 5.1 and Claude Mythos
5.1 System Card,” https://www.anthropic.com/system-cards, 2026. 1
[3] Google DeepMind, “Gemini 3.8 Flash Model Card,” https://deepmind.google/models/model-cards/gemini-3-8-flash/, 2026. 1
[4] DeepSeek-AI, “DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient,” Tech. Rep., 2026. 1
[5] Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan et al., “The llama 3 herd of models,” arXiv preprint arXiv:2407.21783, 2024.
Improvements for AI systems
- Bold Header: Learnable Trust-Boundary Delimiters (LTBD) Implementation
LTBD explicitly encodes trust boundaries by learning a small number of learnable delimiters to distinguish trusted user instructions from untrusted external data, enabling the model to better respect the intended trust hierarchy.
This allows the model to distinguish trusted instruction and potentially adversarial external content
without modifying underlying LLM parameters, offering a lightweight yet explicit structural separation.
- Bold Header: Robustness Against Prompt Injection
The system gains strong robustness by associating learned embeddings with content provenance, meaning the defense is encoded in the learned embeddings associated with the true structural trust boundaries,
making it substantially harder to bypass
even against adaptive attacks like Delimiter-Spoof.
- Bold Header: Utility Preservation During Defense
LTBD achieves a strong security–utility trade-off by optimizing a loss function that ensures preservation of benign behavior, specifically by encouraging the model to preserve the response associated with the original benign task, while explicitly increases the relative likelihood of y+ over the attack-induced response y−.
This means it preserves benign task utility and introduces negligible inference overhead.
Sources
- The Llama 3 Herd of Models
- CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models
- Ignore Previous Prompt: Attack Techniques For Language Models
- Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs