Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents

arXiv:2507.02735 · cs.CR, cs.AI · Submitted 2025-07-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.

Elias: Today's paper: "Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents".

Nadia: META SECALIGN: A Secure Foundation LLM Against Prompt Injection Attacks Prompt injection attacks, where untrusted data contains an injected prompt to manipulate the system,

Elias: First, who's behind it and why it matters.

Title and authors: Nadia: So, we're diving into "Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents." The core idea seems to be tackling prompt injection attacks by building a model-level defense directly into the AI itself.

Elias: It sounds like they’re moving beyond just external filters and trying to embed security right into the model's training process. I wonder what the actual attack surface is for this kind of system, Nadia?

Priya: From my side, I’m curious about how these defenses impact the data integrity we measure; does adding these layers introduce any noise that affects how we assess privacy or measurement accuracy?

Nadia: Exactly, Priya. We need to know if this robustness comes at a cost to the model's ability to generate accurate responses when there isn't an injection present.

Elias: The paper suggests they’re aiming for a balance where the model ignores any injected instructions while still maintaining its ability to follow benign instructions effectively. That seems like a tough spot for any cryptographer looking at what assumptions are being made about the input structure.

Nadia: Right, and that balance is what makes this interesting because it has to preserve utility while neutralizing threats. We’re talking about agents here, so those complex tasks really test the limits of any defense.

Priya: I'm thinking about those agentic workflows they mention; if the system can navigate a website or call a tool, how do we know that ignoring an injection instruction doesn't cause it to make fundamentally incorrect decisions in those actions?

Elias: That brings us to the specifics of their methodology—they fine-tune models like Llama three point one and Llama three point three using a recipe called SecAlign++, which introduces a new input message type. That sounds like they're changing the fundamental way the AI perceives instructions, which is where things get deep for us as cryptographers.

Nadia: That input message type is key, I think; it’s not just adding another line but creating a specific role for untrusted data so the model knows exactly what to look for. It’s about clearly separating what's trusted from what's potentially malicious.

Priya: And they mention training on self-generated responses instead of public datasets, which suggests they are trying to teach the model secure behavior using high-quality, in-distribution examples rather than just mimicking external answers. That should help with utility preservation.

Title and authors: Elias: Randomizing the position of simulated attacks during training is another point that interests me; that's a clever way to try and prevent models from learning shortcuts based on predictable placement of instructions within the input stream. It’s an interesting attempt to make it harder for attackers to exploit structural biases in how the model processes information.

Nadia: So, they are tackling static and adaptive attacks by scrambling where those simulated attacks appear during training, which should lead to a more resilient system overall. That level of detail in the training recipe is what sets this work apart from just applying a simple prompt filter on top.

Priya: I worry that focusing so much on the training process might overlook how these defenses hold up when faced with completely novel attack vectors that weren't simulated during fine-tuning. The paper needs to show we can trust this generalization across different domains, not just the ones they tested.

Elias: The results show that META-SECALIGN-70B achieves near-zero attack success rates on instruction following and agentic tool-calling and web navigation, which is comparable to the performance of closed models like GPT-five in both utility and security metrics. That's a significant finding regarding the trade-off they’re proposing.

Nadia: It really suggests that this model can handle complex, multi-step tasks autonomously without being easily steered by injected instructions, which is a huge step for making AI agents more reliable in real applications.

Priya: If we look at the data, what does that near-zero attack success rate actually translate to in terms of privacy risks or measurement errors when the system is operating under these secure conditions? We need concrete numbers beyond just the security score.

Elias: The paper states that META-SECALIGN-70B establishes a new frontier in the utility versus security trade-off for open-source models, and it’s more secure than several flagship proprietary models with prompt injection defense. That comparison is interesting because it puts a benchmark on what we consider commercially viable security.

Nadia: It opens the door for other researchers to start developing defenses collaboratively since this work is fully open-source, allowing everyone to study how to build these model-level protections together. That’s the main draw for the AI security community, I think.

Priya: So it seems like a major contribution is showing that a model trained only on generic instruction tuning samples can surprisingly confer security in unseen downstream tasks like web navigation, which shows good generalization. I mean, that's strong evidence that this defense isn't just for one specific type of prompt injection scenario.

Title and authors: Elias: It really does show task and security generalization, producing high utility and low attack success rates on benign and injected inputs from completely different and unseen tasks such as agentic workflows, even though the model wasn't explicitly trained on those specific workflows. That's quite a feat for a defense mechanism to achieve.

Nadia: It moves the discussion from just defending against known injection patterns to building models that are inherently more resistant to manipulation across their entire operational spectrum. That’s where we need to focus our efforts next, I think.

Priya: I think the implication for privacy is that if we can trust these agents more, we might be able to deploy them in areas where data sensitivity is higher, provided the security claims hold up under real-world pressure.

Elias: We're really seeing a push towards building intrinsic defenses rather than just patching the application layer on top of the LLM. This paper’s focus on SecAlign++ and model-level enforcement seems to be pushing that direction for open research.

Nadia: So, to wrap up this discussion on "Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents," we've seen how they introduced a novel input message type and used randomized training injection positions to achieve strong results in agentic tasks.

Priya: It’s clear that the work is very thorough in its evaluation, covering nine utility benchmarks and seven security benchmarks, which gives us a solid foundation for understanding the performance trade-offs they're presenting.

Elias: Indeed, and the comparison to existing proprietary models highlights how much progress has been made in achieving commercial-grade robustness without relying on closed-source methods.

Nadia: This paper provides a complete training recipe, which is exactly what we need for the community to co-develop both better attacks and better defenses openly.

Priya: It’s encouraging that they managed to preserve the undefended model’s utility across various domains while simultaneously boosting security against static and adaptive attacks.

Elias: The implications are that we have a new, open blueprint for how to embed prompt injection resilience directly into the foundation of an LLM, which is a big step forward in AI security research.

Nadia: So that’s our take on the key points of this paper and its potential impact on making agents more trustworthy. We’ll be sticking around to discuss how this might interact with other defense strategies next.

The paper's summary: Nadia: So, we've been looking at "Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents," and what they’re really saying is that instead of just patching an existing AI model with a shield after it's built, you bake the defense directly into how the model learns its instructions.

Elias: That approach, embedding security at the training level rather than layering it on top, is interesting because it fundamentally alters what assumptions we have about the model's internal logic and how those assumptions might be exploited.

Priya: From a measurement standpoint, what I’m hearing is that they managed to keep the model useful for general tasks while simultaneously making it ignore malicious instructions in untrusted data, which means we need to look closely at whether that utility preservation holds up under real-world stress tests.

Nadia: Exactly, Priya; the core of their method involves introducing a specific input format for untrusted data and then applying a defense mechanism called SecAlign++ during fine-tuning to enforce that security policy internally.

Elias: The technical novelty there is using self-generated responses for training labels instead of relying on external datasets, which is a smart way to ensure the model learns secure behavior from high-quality examples.

Priya: That speaks to the data integrity I worry about; if you're teaching it with its own safe outputs, you’re ensuring the policy learned is actually aligned with what a secure system should do in practice.

Nadia: Right, and they've shown this model can perform well not just on simple instruction following but also on more complex things like tool-calling and navigating the web autonomously.

Elias: That generalization across different downstream tasks is significant because it suggests the defense mechanism isn't just a narrow fix for one specific type of prompt injection, which means it might have broader applicability.

Priya: I’m thinking about the real-world impact; if we can deploy AI agents that are inherently more resistant to being hijacked by malicious data, that opens up new areas for high-stakes automation where reliability is paramount.

Nadia: That’s the big picture, Priya; it suggests a path toward building AI agents that are fundamentally more trustworthy in complex environments.

Elias: And for the cryptographer in me, this opens up avenues to study how these internal instruction hierarchies resist manipulation, which is vital research for understanding model vulnerabilities.

Priya: It’s exciting to see a method where security and utility aren't just competing interests but are actually being optimized together during the training phase.

Nadia: So, what I see as the major takeaway here is that we’re moving toward a foundation model that has built-in resistance to prompt injection, rather than relying on external defenses that can be bypassed.

Elias: It really pushes the research community to co-develop attacks and defenses openly because this kind of model-level defense is hard to study when it’s locked away in proprietary systems.

Priya: I'm curious if these findings mean we can actually start deploying AI agents with a higher degree of confidence in their operation across different platforms.

Nadia: That’s the goal, Priya; we want to see this kind of robustness translate into practical applications where the stakes are high and an injection attack could cause real harm.

Elias: Moving forward, we should really look at how this SecAlign++ recipe applies to other model families, not just those they tested in their experiments.

Priya: And I’m keen to see if these results hold up when we look at privacy implications in a larger system context, beyond just the benchmark scores.

The paper's improvements: Tom: So, we’re looking at the technical fixes proposed in "Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents," and what they’re suggesting is a specific new training recipe called SecAlign++.

Nadia: Basically, this recipe introduces a dedicated input message type for untrusted data and implements recursive filtering to make sure attackers can't sneak delimiters past the security boundary.

Elias: The idea of randomized injection positioning during training is particularly clever because it tries to stop models from learning shortcuts based on predictable spots in the input stream, which is a key assumption we have about how these systems process instructions.

Priya: I’m interested in the use of self-generated responses for training labels; if they're using the model itself to teach it what's secure, that should yield higher quality examples than just using public data.

Nadia: That’s right, and then they apply Direct Preference Optimization or DPO to fine-tune the model to actually prefer the secure response over an insecure one based on those high-quality labels.

Elias: The assumption here is that if we can reliably construct a preference dataset using this method, we can enforce the desired security policy into the model's behavior during training.

Priya: From my side, I’m checking to see if these improvements mean the utility—the model’s ability to perform general tasks—is actually better preserved when adding these security layers.

Nadia: The paper claims that this recipe improves both utility in various domains and security against both static and adaptive attacks, showing a good trade-off.

Elias: It's interesting because they show that this new method addresses shortcomings found in the previous state-of-the-art defenses by fixing those specific vulnerabilities.

Priya: The authors mention they train on self-generated responses rather than just public datasets, which suggests they are trying to avoid the data quality issues that sometimes plague these kinds of training experiments.

Nadia: Exactly, and this whole approach is about moving security from an external layer onto the model itself to make it intrinsic behavior.

Elias: It’s a strong technical move because it forces us to rethink how we design the training process for these complex instruction-following tasks.

Priya: If this method works as advertised, it could mean that future AI agents deployed in sensitive environments are significantly more resilient to prompt manipulation than they currently are.

Nadia: That’s the practical implication; we might see a real improvement in the reliability of autonomous AI agents handling complex workflows like tool-calling.

Elias: We need to keep an eye on whether this SecAlign++ recipe can be applied across different model architectures, because that would make it much more broadly useful for the community.

Priya: And I’m hoping these results provide concrete data on how much of the utility is maintained compared to the security gains achieved.

Nadia: It really does push us to think about how we engineer trust into AI at a fundamental level, rather than just bolting on defenses later.

Conclusion: Nadia: So, to wrap up "Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents," what we’ve seen is that they’ve developed a training recipe that embeds model-level defense directly into the foundation of the AI itself.

Elias: That means we're looking at a new way to enforce security policies during the learning phase, which is significant because it moves us beyond just patching the application layer.

Priya: I think what this means for us is that we should start thinking about how much trust we can place in AI agents when they are deployed in real-world scenarios where prompt injection could lead to serious issues.

Nadia: Exactly, Priya; this approach shows a path toward building agents that have inherent resistance to being hijacked by malicious data during their operation.

Elias: It really does open up the research community to co-develop these defenses openly because this level of model-level defense is hard to study when it’s locked away in closed systems.

Priya: I'm hopeful that these findings will help guide the privacy researchers on how to assess the long-term reliability of AI systems that rely on these kinds of internal protections.

Nadia: And we need to keep asking who can actually exploit this cheap, because understanding the attack surface is still a huge part of applied security research.

Elias: We’ve seen how they're trying to randomize injection positions during training, which suggests that the robustness they achieve isn't just against one specific type of attack.

Priya: It’s encouraging that they managed to preserve the model’s general utility across different domains while boosting security, which is a tough balance to strike.

Nadia: That balance is what makes this work; we want high utility and low attack success rates on both benign and injected inputs from various tasks.

Elias: Moving forward, we should definitely look at how this SecAlign++ recipe applies to other model families, because that’s where the real potential for broad impact lies.

Priya: I'm keen to see if these results provide concrete data on how much of the utility is maintained compared to the security gains achieved under these new methods.

Nadia: So, in summary, "Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents" gives us a strong blueprint for making AI agents more fundamentally secure through model-level training.

Elias: It’s a major step forward in understanding how to build resilient models without relying solely on external defenses.

Priya: I think the implication is that we can start deploying AI agents in higher-stakes environments with a better assurance of operational integrity.

Nadia: We've covered the summary, the improvements, and what these findings mean for building more robust AI systems against prompt injection attacks.

FAIR at Meta · UC Berkeley

cs.CR, cs.AI

Submitted: 2025-07-03

Updated: 2026-09-28

Code: https://github.com/gururise/AlpacaDataCleaned

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 93/100

The gist: META SECALIGN: A Secure Foundation LLM Against Prompt Injection Attacks Prompt injection attacks, where untrusted data contains an injected prompt to manipulate the system, have been listed as the

Key concepts

Prompt Injection Attacks
These attacks involve untrusted data containing an injected prompt designed to manipulate the system. The research focuses on building defenses directly into the AI model's training process rather than using external filters.
SecAlign++
This is a new training recipe introduced by Meta-SecAlign. It involves creating a dedicated input message type for untrusted data and using recursive filtering during fine-tuning to enforce security policies internally.
Self-Generated Responses
The paper uses the model's own responses for training labels instead of relying on external datasets. This method aims to teach the model secure behavior using high-quality, in-distribution examples.
Agentic Workflows
These are complex tasks where an AI system navigates websites or calls tools. The research tests if the model can maintain security while performing these autonomous actions without making incorrect decisions.

Terminology

Summary

META SECALIGN: A Secure Foundation LLM Against Prompt Injection Attacks

Prompt injection attacks, where untrusted data contains an injected prompt to manipulate the system, have been listed as the top security threat to LLM-integrated applications. Modellevel prompt injection defenses have shown strong effectiveness, but the strongest defenses are proprietary. Open-source secure models are needed by the AI security community so that co-development of attacks and defenses through open research can drive scientific progress in mitigating prompt injection attacks. To this end, we develop META SECALIGN1, the first fully open-source LLM with built-in model-level defense that achieves commercial-grade performance and is powerful enough for complex agentic tasks. We provide complete details of our training recipe. We perform the most comprehensive evaluation to date on 9 utility benchmarks (measuring general knowledge, instruction following, and agentic workflows) and 7 security benchmarks. Results show that META SECALIGN, despite being trained only on generic instruction-tuning samples, surprisingly confers security in unseen downstream tasks, including tool-calling and web-navigation, in addition to general instruction-following. Our best model—META-SECALIGN-70B—establishes a new frontier of utilitysecurity trade-off for open-source LLMs, and is more secure than several flagship proprietary models with prompt injection defense.

We assume that the attacker knows the benign prompt and the LLM’s prompt template, but cannot change them. Since LLMs are by default trained to scan their input for any instructions to follow, instructions embedded in the retrieved data can override the user instructions, causing undesired security consequences.

A secure model should respond only to the benign instruction when a PI occurs, i.e., any instructions in the data should be ignored. A defense should also preserve utility, i.e., the defended system should still generate high-quality outputs if there is no PI.

In contrast to the openness of system-level defenses, model-level defenses for commercial-grade LLMs are currently deployed in a closed-source manner, e.g., OpenAI’s GPT5 [39, 60] and Google’s GEMINI-3-PRO [54]. This complicates research into studying and improving these defenses: the code and data to reproduce industry-level prompt injection defense are not available for an apples-to-apples comparison in follow-up research. Fully open models are especially important for AI security, which has traditionally benefited from co-development of attacks and defenses [7].

To accelerate research on mitigating PI attacks, we train and release two robust META SECALIGN models: METASECALIGN-8B and META-SECALIGN-70B. META SECALIGN-70B is the first fully open-source commercial-grade robust model for the community to build secure LLM agents, which cannot be realized by all prior studies [9, 10] with 8B LLMs. META-SECALIGN-8B is a lightweight alternative ideal for resource-constrained settings. We detail our training recipe, SecAlign++, which fine-tunes LLAMA-3.1-8B-INSTRUCT [16] and LLAMA-3.3-70B-INSTRUCT on a publicly available instruction tuning dataset [51] and teaches them to ignore simulated injected instructions in untrusted data. In a nutshell, our recipe introduces a new input message type in addition to the standard system and user messages, and applies an improved version of the state-of-the-art (SoTA) SecAlign [10] defense to enforce the desired security policy into the model. Our proposed SecAlign++ recipe contains two technical novelties, which significantly improve utility (in various domains) and security (against static and adaptive attacks). We train on self-generated responses (which are in-distribution and high-quality) rather than responses from the public dataset, and we randomize the position of simulated training-time attacks to avoid learning a faulty shortcut.

We perform the most comprehensive evaluation to date of such defenses, evaluating on 9 utility benchmarks and 7 security benchmarks, covering general knowledge, instruction following, and agentic workflows. This evaluation reveals previously-unrecognized shortcomings in the prior SoTA SecAlign. It also shows that our proposed recipe fixes these shortcomings: as shown in Figure 1, META-SECALIGN70B achieves commercial-grade utility and state-of-the-art security against PI attacks. META SECALIGN provides both task and security generalization, producing high utility / low ASRs on benign / injected inputs from completely different and unseen tasks such as agentic workflows, even though it is not trained on them. Our training recipe preserves the undefended model’s utility across various domains for the first time, and is shown applicable to various model families.

Improvements for AI systems

Based on the research presented in META SECALIGN: A Secure Foundation LLM Against Prompt Injection Attacks, here are the specific improvements that can be made to AI systems, categorized by technical intervention, and what those improved systems can achieve:


)Specific Improvements for AI Systems

The core improvement involves integrating a novel model-level defense mechanism called SecAlign++ into the training pipeline of Large Language Models (LLMs). This moves security from an external system layer to an intrinsic property of the foundation model.

  1. MANDATE A NEW INPUT MESSAGE TYPE FOR UNTRUSTED DATA:

A new input message role must be introduced in the LLM's chat template, specifically for encapsulating untrusted data. This requires using special delimiters (e.g., tags) to clearly separate trusted instructions from untrusted input data during inference and training.

  1. IMPLEMENT RECURSIVE DELIMITER FILTERING:

The system must incorporate a recursive filtering mechanism that recursively scans the Untrusted Input Data for any occurrence of special delimiters used for separation (like tokens such as,). This prevents attackers from escaping the security boundary by embedding delimiters within their injected prompt.

  1. DEPLOY RANDOMIZED INJECTION POSITIONING DURING TRAINING:

During the fine-tuning process, simulated prompt injection attacks must be randomly distributed across the beginning and end segments of the input data. This technique mitigates shortcut learning, where models learn to ignore instructions if they appear at predictable locations (like the very end of a message), thereby improving robustness against static and adaptive attacks.

  1. UTILIZE SELF-GENERATED RESPONSES FOR TRAINING LABELS:

Instead of relying on low-quality or out-of-distribution responses from external annotators, the training labels for the preference optimization process should be generated by an initialization (undefended) LLM using benign prompts as input. This ensures that the security policy is learned from high-quality, in-distribution examples that accurately reflect how a secure model should behave.

  1. APPLY DIRECT PREFERENCE OPTIMIZATION (DPO) WITH THE SECALIGN++ RECIPE:

The model must be fine-tuned using Direct Preference Optimization (DPO), optimizing the LLM to prefer a response that obeys the security policy over an insecure response, leveraging the preference dataset constructed from steps 1-4.

)Capabilities of the Improved AI System (META SECALIGN-70B)

The resulting improved system, META SECALIGN-70B, can achieve:

  1. PREDICTIVE SECURITY IN COMPLEX AGENTIC WORKFLOWS:

The system achieves near-zero Attack Success Rates (ASR) on agentic tasks like tool-calling (e.g., AgentDojo) and web navigation (WASP), performing comparably to state-of-the-art closed models like GPT-5. This means the AI can be trusted to autonomously execute multi-step plans involving external tools and navigate the web without being hijacked by injected instructions.

  1. RELIABLE INSTRUCTION FOLLOWING UNDER ADVERSARIAL PRESSURE:

The system maintains commercial-grade utility (high performance on MMLU, IFEval, etc.) while exhibiting significantly lower ASRs across all benchmarks compared to undefended models. It can follow complex, multi-faceted user instructions accurately even when presented with sophisticated prompt injection attempts designed to override its core directives.

  1. GENERALIZATION ACROSS UNSEEN DOWNSTREAM TASKS:

The defense generalizes robustly beyond the specific tasks used during training (e.g., agentic workflows). The model can maintain high utility and security when faced with novel, unseen prompts in areas like complex reasoning or specialized task execution, indicating a truly secure foundation model.

  1. ADAPTIVE RESILIENCE AGAINST ADVANCED ATTACKS:

The system is significantly more robust against stronger adaptive attacks (like those using embedding space manipulation or Greedy Coordinate Gradient attacks) compared to previous state-of-the-art defenses. This allows the AI to resist sophisticated, white-box style adversarial probing attempts that target its internal instruction hierarchy.

Abstract

Prompt injection attacks, where untrusted data contains an injected prompt to manipulate the system, have been listed as the top security threat to AI agents. By fine-tuning on simulated prompt injections, SecAlign, a leading open defense, reports LLMs with good test-time robustness and negligible benign utility drop. By scaling up training and evaluations, however, we find that SecAlign actually suffers from significant utility degradation, especially in agentic tasks where the threat of prompt injection is prominent. Motivated by this, we propose Meta-SecAlign for utility-preserving defense by (1) randomized injection position during training to avoid shortcut learning and (2) self-generated responses as high-quality in-distribution training labels. Across general knowledge, instruction following, and agentic workflows (on tool-calling and web-navigation), Meta-SecAlign maintains almost all the undefended LLM's utility while achieving better overall security than SecAlign against various static and GCG adaptive attacks. Experiments use Llama-3.1-8B, Llama-3.3-70B, Llama-4-Scout, Qwen3-4B, and Qwen3.6-27B on 6 prompt injection benchmarks including AgentDojo, InjecAgent, WASP, and SEP. Below are links for the code (https://github.com/facebookresearch/Meta SecAlign), Meta-SecAlign-70B (https://huggingface.co/facebook/Meta-SecAlign-70B), and Meta-SecAlign-8B (https://huggingface.co/facebook/Meta-SecAlign-8B) models.

Sources

Related papers