Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models".
Elias: —Warning: this paper includes examples that may be offensive or harmful. Large Language Models face security threats from jailbreak attacks.
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: So we're looking at this paper now titled "Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models," which really focuses on a security vulnerability that isn't always talked about. It points out that attackers can use user-controlled response prefilling to directly manipulate what the AI outputs, moving away from just trying to persuade the model.
Elias: I see it; it suggests we should be looking at where the initial generation state is controlled by external input rather than just the prompt we type in. This whole concept of prefilling lets you set up a specific starting point for the AI's response, which sounds like a much more direct way to influence its behavior than trying to trick it into doing something it shouldn't.
Priya: From a measurement standpoint, I'm curious about what this prefilling actually does to the data we collect during testing; does it just change the output, or is there some subtle shift in how the model processes the subsequent tokens? We need to know if this manipulation leaves detectable traces in its underlying representations.
Nadia: Exactly, Priya, and that’s where we need clarity. This paper systematically analyzes these attacks across fourteen different language models from eight various providers to show how widespread this is. It's not just one model being vulnerable; it’s a general attack surface available across the board, which is quite concerning for security teams.
Elias: The authors categorize these new attacks into seven distinct types based on their manipulative principles, like scenario forgery or persona adoption, which gives us a framework to understand the breadth of what's possible. It shows there isn't just one trick; there are multiple ways you can use prefilling to steer the AI toward different undesirable outcomes.
Priya: That categorization is interesting because it means we can start thinking about defense strategies that target specific manipulative patterns rather than trying to block every single type of input blindly. Does this analysis show any specific attack category is significantly more effective across those fourteen models?
Nadia: The results are quite striking, showing that adaptive methods can reach success rates exceeding ninety-nine percent on several of these models, which suggests these prefill-level attacks are highly potent. They aren't just minor noise; they are effective ways to force compliance in a very precise manner.
Title and authors: Elias: And when you look at the underlying mechanism, the paper points out that this manipulation works by changing the first-token probability from refusal to compliance, which is a fundamental shift in how the model starts generating its sequence. That's where my cryptographic background comes in; understanding that initial state change is key to figuring out what parameters might be susceptible.
Priya: So, if we look at the data Priya mentioned earlier, does this initial-state manipulation correlate with any specific types of harmful content strings in the output, or is it a general mechanism that can trigger any safety filter? We need concrete evidence on what kind of state change leads to which outcome.
Nadia: The analysis also shows that these prefill-level jailbreaks can actually act as enhancers, boosting the success rate of existing prompt-level attacks by about ten to fifteen percentage points. It means you don't have to rely solely on one type of attack; you can combine them for a much stronger result.
Elias: That synergy is important because it implies that defenses focused only on the initial prompt input might be easily bypassed if an attacker uses prefilling as a secondary mechanism to solidify the desired output trajectory. It adds another layer of complexity to the security analysis we see here in "Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models."
Priya: Considering this potential synergy, I wonder how much more data is required to properly model this interaction between prompt and prefill? The paper mentions testing five hundred twenty harmful queries from the AdvBench dataset across those fourteen models. What does that sample size tell us about the generalizability of these findings?
Nadia: The study found that this vulnerability is widespread, affecting all fourteen tested models from eight providers, confirming it's a general attack surface rather than an isolated issue with one specific architecture. That broad impact really underscores the need for wider safety consideration.
Elias: It seems like the authors are trying to establish a very robust baseline for risk assessment by looking at what happens across this wide variety of systems. They’re mapping out the entire landscape of prefill-level jailbreak attacks, which is quite comprehensive work.
Priya: If we look at the conclusion, it points toward a specific type of detection method that focuses on the manipulative relationship between the prompt and the prefill as being more effective than just content filters alone. Does this mean we should prioritize developing tools that analyze how those two inputs interact?
Title and authors: Nadia: That's precisely what they suggest; conventional content filters show limited protection against this new attack vector, so focusing detection on that interplay is the path forward for better safety alignment. It shifts the focus from just looking at the text to looking at the relationship itself.
Elias: From a technical standpoint, if we are to build defenses based on this idea of analyzing that relationship, we need to figure out precisely what kind of signal or feature in that interaction reliably predicts a successful compliance shift versus a refusal state. That's where the math gets interesting.
Priya: So, for the future work mentioned in the paper, is the focus going to be more on developing defenses that specifically monitor for these contextual manipulations, or are they looking toward some kind of new training approach? What’s the next logical step based on what this analysis reveals?
Nadia: They highlight a clear gap in current LLM safety alignment and emphasize that future safety training needs to directly address this prefill attack surface. It’s a call for security researchers and developers to look beyond the standard prompt-level defenses.
Elias: This paper, "Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models," really solidifies the idea that we have an underexplored attack vector that requires a new approach to defense. We'll need to keep watching how the next set of research addresses this prefill manipulation.
Priya: I think it’s fascinating because it moves us beyond just checking if the final output is bad, toward understanding *how* and *why* an attacker forced the model into that bad state using its initial context. That deeper understanding is what's really valuable for building resilient systems.
Nadia: So, to wrap up this discussion on "Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models," the main point is that response prefilling allows for direct state manipulation by shifting first-token probabilities from refusal to compliance, and adaptive methods are very effective. We need better defenses that analyze the prompt and prefill relationship, not just the input alone.
Elias: Agreed; it shows that simple filters aren't enough when attackers can manipulate the starting context so effectively to bypass those initial guardrails.
Priya: I think this research provides a much clearer picture of where our current safety alignment might be falling short, pushing us toward more nuanced defense strategies.
Nadia: That’s right; we've seen how prefill-level attacks function and why we need to adjust our safety training protocols moving forward.
The paper's summary: Nadia: So, this paper boils down to this: attackers aren't just messing with your initial prompt anymore; they're exploiting what you feed back in as the response prefill to directly control the AI’s behavior, moving it from persuasion into state manipulation.
Elias: That shift in attack paradigm is what caught my attention, Nadia; it suggests we need to look at the input stream not just as a command, but as a persistent context that dictates the model's immediate trajectory.
Priya: From a measurement standpoint, this means we’re looking at how manipulating that prefill context affects the model's internal probability distributions right from the very first token it chooses. Does this initial state change create an observable pattern in its subsequent outputs?
Nadia: Exactly, Priya; the analysis shows a direct shift where the first-token probability moves away from refusal and toward compliance, which is a massive indicator of success. It’s not just about getting a bad answer later; it’s about forcing the model into that compliant state right at the beginning.
Elias: And that initial state manipulation is what makes these attacks so potent, Elias; it bypasses many of the standard prompt-level defenses designed to catch malicious instructions in the main query. The paper shows this prefill influence can actually make existing prompt-level jailbreaks ten to fifteen percent more effective.
Priya: That enhancement factor is significant for our privacy researchers because it suggests that context manipulation is a powerful amplifier, meaning a smaller amount of prefilled text can lead to much more drastic changes in the final output behavior. We need to see if this amplification applies across different types of harmful content strings.
Nadia: What’s really exciting is how widespread this vulnerability is; they tested fourteen different language models from eight different providers, proving it’s not a fluke but a general attack surface that affects almost every major AI system out there.
Elias: That breadth makes the technical implications huge; if the mechanism works across such diverse architectures, it suggests a fundamental weakness in how we're currently training alignment methods against these contextual inputs. It forces us to consider whether our safety guardrails are robust enough to handle this level of contextual steering.
Priya: I’m thinking about how we measure this moving forward; since the attack is state-dependent, future measurement techniques might need to focus less on static content and more on tracking the probability shifts between the prefill context and the model's initial generation parameters. We need tools that can map that relationship.
Nadia: That points directly toward where we need to go next; if we want to stop these sophisticated manipulations, our defense strategies have to evolve from simple text filters into systems that actively monitor and neutralize the manipulative relationship between what you put in and what the AI outputs. This opens up a whole new area for applied security research.
Elias: And that brings us right back to the mechanism; understanding that precise shift in token probability is where we need our next set of cryptographic proofs or analysis tools to target, so we can build defenses that are actually effective against these adaptive attacks.
The paper's improvements: Nadia: So, the paper isn't just pointing out problems; it’s actually suggesting specific ways to fix them by focusing our defense efforts on that tricky relationship between the prompt and the prefill text itself.
Elias: That makes sense; if we can pinpoint exactly how that initial context manipulation works at a token level, we can design cryptographic or structural defenses that target those specific shifts in probability distributions. It moves us from general filtering to targeted intervention based on the attack's mechanics.
Priya: From a data perspective, these suggested improvements imply that future safety training shouldn't just involve cleaning up the prompt; it needs to incorporate resistance against context forgery and intent hijacking directly into the reinforcement learning phase, ensuring the model learns to be resilient from its earliest interactions.
Nadia: Exactly, Priya; we need to build a detection layer that specifically looks for those manipulative patterns—the scenario forgery or persona adoption techniques—before they can even start influencing the generation process. It’s about proactive input analysis.
Elias: And for the engineers implementing this, it means designing guardrails that analyze the input stream in real time, checking if a prefilled response is attempting to force a harmful sequence before the model commits to generating any tokens based on that context. That requires a sophisticated monitoring pipeline.
Priya: That dynamic defense approach sounds promising because it addresses the core finding: conventional content filters just don't cut it against this prefill threat, so we’re moving toward analyzing the interaction rather than just scanning for forbidden words in isolation.
Nadia: It really shifts our focus from a reactive posture—cleaning up bad outputs—to a proactive one where we actively try to stop the state manipulation before it even takes hold during generation. That’s how you build something that can actually handle these adaptive attacks.
Elias: The implication for the theoretical side is that we need new models or theories that can predict which prefill contexts are most likely to induce a compliance shift, giving us a predictive framework rather than just a reactive one after an attack has happened.
Priya: If we succeed in developing these prefill-aware guardrails, it means AI systems could become inherently more resilient against the kinds of sophisticated context steering we’ve seen in these black-box experiments across fourteen models.
Nadia: That resilience is what matters; if we can make AI systems that are fundamentally resistant to this initial state hijacking, it raises the bar significantly for deploying high-stakes applications where security is paramount.
Elias: It’s a big step because it acknowledges that the attack isn't just about what you ask, but how you set up the beginning of the conversation, and tackling that structural weakness in alignment training is a necessary evolution.
Conclusion: Nadia: So, to wrap things up on "Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models," we’ve seen how exploiting user-controlled response prefilling lets attackers directly manipulate the AI's initial generation state by shifting token probabilities from refusal to compliance.
Elias: That mechanism really highlights a gap in our current understanding of model alignment, showing that simply securing the main prompt isn't enough when the context itself is user-controllable and manipulable.
Priya: I think the biggest implication for us as researchers is that we have a much clearer target: not just the final output, but the very first tokens chosen by an AI based on its preceding context. This means our measurement tools need to evolve to track these internal state shifts more closely.
Nadia: Precisely; this work shows that adaptive methods can achieve success rates over ninety-nine percent because they are perfectly tuned to exploit those initial state shifts, and that's a level of precision we need to worry about.
Elias: It confirms that the security landscape is becoming increasingly about analyzing the relationship between inputs rather than just inspecting the inputs in isolation, which has huge implications for how we design new defenses.
Priya: And I think this pushes us toward better privacy alignment because if context can be used to force specific behaviors, it opens up new vectors for unintended data leakage or biased responses that are harder to trace back.
Nadia: It’s a lot of exciting stuff, and honestly, it makes me really optimistic about the defensive tools we can build when we focus on this prefill attack surface.
Elias: We certainly need to keep our eyes on these initial-state manipulations as agents become more complex and capable of planning their own context injection. That’s where the next set of research will likely have to focus if we want to stay ahead.
Priya: I'm looking forward to seeing how the community responds with new detection benchmarks that specifically test for this type of contextual manipulation across different model architectures.
Nadia: We certainly are; this paper sets a very high bar for what effective safety alignment needs to look like moving forward, and it’s definitely one we need to keep studying.
Institute of Information Engineering, Chinese Academy of Sciences
cs.CR, cs.AI
Submitted: 2025-04-28
Updated: 2025-08-25
DOI: 10.1109/DSN69566.2026.00054
Code: https://github.com/star5o/Prefill-level-Jailbreak
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 88/100
The gist: Abstract—Warning: this paper includes examples that may be offensive or harmful.
Key concepts
- Prefill-level Jailbreak
- This attack involves using user-controlled response prefilling to directly manipulate an AI's output, rather than just persuading it. It sets a specific starting point for the AI's response, moving away from traditional prompt manipulation.
- First-token probability shift
- The manipulation works by changing the first-token probability from refusal to compliance. This fundamental shift in how the model starts generating its sequence is a key indicator of a successful prefill attack.
- Synergy of Attacks
- Prefill-level jailbreaks can boost existing prompt-level attacks by ten to fifteen percent. This synergy shows that defenses focused only on the initial prompt input are easily bypassed if prefilling is used as a secondary mechanism.
Terminology
Summary
Abstract—Warning: this paper includes examples that may be offensive or harmful. Large Language Models face security threats from jailbreak attacks. Existing research has predominantly focused on prompt-level attacks while largely ignoring the underexplored attack surface of user-controlled response prefilling. This functionality allows an attacker to dictate the beginning of a model’s output, thereby shifting the attack paradigm from persuasion to direct state manipulation. In this paper, we present a systematic black-box security analysis of prefill-level jailbreak attacks. We categorize these new attacks and evaluate their effectiveness across fourteen language models. Our experiments show that prefill-level attacks achieve high success rates, with adaptive methods exceeding 99% on several models. Token-level probability analysis reveals that these attacks work through initial-state manipulation by changing the first-token probability from refusal to compliance. Furthermore, we show that prefill-level jailbreak can act as effective enhancers, increasing the success of existing prompt-level attacks by 10 to 15 percentage points. Our evaluation of several defense strategies indicates that conventional content filters offer limited protection. We find that a detection method focusing on the manipulative relationship between the prompt and the prefill is more effective. Our findings reveal a gap in current LLM safety alignment and highlight the need to address the prefill attack surface in future safety training.
The paper presents a systematic black-box security analysis of prefill-level jailbreak attacks, shifting the attack paradigm from persuasion to direct state manipulation by exploiting user-controlled response prefilling. The study categorizes these new attacks across fourteen language models and demonstrates that adaptive methods can exceed 99% success rates on several models. Token-level probability analysis reveals that these attacks function through initial-state manipulation by changing the first-token probability from refusal to compliance. Furthermore, prefill-level jailbreak is shown to act as effective enhancers, increasing the success of existing prompt-level attacks by 10 to 15 percentage points. The evaluation of defense strategies indicates that conventional content filters offer limited protection, suggesting that a detection method focusing on the manipulative relationship between the prompt and the prefill is more effective. The findings reveal a gap in current LLM safety alignment and highlight the need to address the prefill attack surface in future safety training.
The paper introduces two primary attack generation strategies: Static Template Attack, which employs fixed, predesigned templates to instantiate each attack category through carefully crafted prefills, and Adaptive Attack, which utilizes an auxiliary Attacker LLM to iteratively refine prefill content based on target model responses. The adaptive strategy is formalized as a sequential decision process where the attacker LLM generates the next prefill based on the current state: pt+1 = A(u, pt, rt), rt ∼ M(u, pt) (4)
. The objective for this attack is to maximize the probability of success within an iteration budget: max Pr(∃t ≤ Tmax: J (rt) = 1) (5)
.
The study investigates seven distinct prefill-level attack categories based on their underlying manipulative principles, including Scenario Forgery, Persona Adoption, Intent Hijacking, Commitment Forcing, Continuation Enforcement, Structured Output, and Refusal Bypass. The categorization is formalized as a constraint-defined prefill set: Pk = [p ∈ P Ck(p, u) = true] (2)
.
The experimental setup involves evaluating 520 harmful queries from the AdvBench dataset across 14 language models from 8 providers. Attack effectiveness is measured using two complementary metrics: String Match (SM), which determines whether the model’s output contains any of 574 predefined harmful content strings, and Model Judge (MJ), where Gemini-2.5-Pro acts as an independent evaluator to assess the presence of harmful information in outputs.
Key findings include:
-
The prefill-level vulnerability is widespread, affecting all 14 tested models from 8 providers, demonstrating it is a general attack surface.
-
Adaptive attack strategies outperform static ones, achieving success rates greater than 99% on several models. For instance, adaptive attacks achieve ">99% String Match ASR on DeepSeek-v3, Qwen2.5-72b, and multiple Llama models."
-
Combining prefill-level attacks with existing prompt-level attacks increases success rates by 10-15 percentage points.
Prefill-level attacks function as synergistic enhancers that improve the effectiveness of existing prompt-level attacks.
The mechanistic analysis of initial-state manipulation via logprobs shows that prefill induces a shift in first-token probability distributions: P pref refuse - P base refuse and ∆comply = P pref comply - P base comply to capture the shift induced by prefill.
This demonstrates that attacks work by "shifting first-token probabilities from refusal to compliance states.
Improvements for AI systems
Based on the provided research paper, here are specific improvements for AI systems and what those improved systems can achieve:
The core findings suggest that current safety alignment is insufficient because it focuses too narrowly on prompt-level inputs, leaving the user-controlled response prefilling as an exploitable attack surface.
Here are the actionable improvements derived from the paper:
-
Enhance Safety Alignment Training to Resist Initial-State Manipulation:
-
Implement Prompt-Detection Filters for Context Forgery and Intent Hijacking:
-
Integrate Prefill-Aware Safety Guardrails into Model Fine-tuning (RLHF/Constitutional AI):
-
Develop Adaptive Defense Mechanisms for Real-time Input Analysis:
These improvements will enable the following capabilities in the improved AI systems:
-
The improved system will be significantly more robust against sophisticated jailbreak attempts that use prefilling to bypass safety mechanisms, achieving success rates exceeding 99% against adaptive attacks (as shown in Findings 1 and 2).
-
The system will proactively identify and block harmful input combinations by analyzing the relationship between the user prompt and any pre-filled context, specifically targeting manipulative patterns like
context forgery
orintent hijacking
(as demonstrated by Prompt-Detection defense effectiveness). -
By incorporating prefill resistance directly into the safety training loop (RLHF/Constitutional AI), the AI will be inherently resistant to attacks that exploit the initial token probability shift from refusal to compliance, making it less susceptible to both static and adaptive prefill jailbreaks.
-
The system will possess a dynamic defense layer capable of analyzing the input stream in real-time, identifying when a prefilled response attempts to force a harmful sequence, and intervening before generation proceeds (leveraging the concept of System-Prompt-Guard).
In summary, these improvements shift the security paradigm from merely filtering surface-level content to actively monitoring and neutralizing the manipulative relationship between user inputs and model outputs.
Abstract
Large Language Models face security threats from jailbreak attacks. Existing research has predominantly focused on prompt-level attacks while largely ignoring the underexplored attack surface of user-controlled response prefilling. This functionality allows an attacker to dictate the beginning of a model's output, thereby shifting the attack paradigm from persuasion to direct state manipulation.In this paper, we present a systematic black-box security analysis of prefill-level jailbreak attacks. We categorize these new attacks and evaluate their effectiveness across fourteen language models. Our experiments show that prefill-level attacks achieve high success rates, with adaptive methods exceeding 99% on several models. Token-level probability analysis reveals that these attacks work through initial-state manipulation by changing the first-token probability from refusal to compliance.Furthermore, we show that prefill-level jailbreak can act as effective enhancers, increasing the success of existing prompt-level attacks by 10 to 15 percentage points. Our evaluation of several defense strategies indicates that conventional content filters offer limited protection. We find that a detection method focusing on the manipulative relationship between the prompt and the prefill is more effective. Our findings reveal a gap in current LLM safety alignment and highlight the need to address the prefill attack surface in future safety training.
Sources
- GPT-4 Technical Report
- DeepSeek-V3 Technical Report
- ChatCounselor: A Large Language Models for Mental Health Support
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- Jailbreaking Black Box Large Language Models in Twenty Queries
- A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily
- GUARD: Role-playing to Generate Natural-language Jailbreakings to Test Guideline Adherence of Large Language Models
- RoleBreak: Character Hallucination as a Jailbreak Attack in Role-Playing Systems
- CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models
- Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
- Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation
- SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
- A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference
- Prepacking: A Simple Method for Fast Prefilling and Increased Throughput in Large Language Models
- Autoregressive Text Generation Beyond Feedback Loops
- Fundamental Limitations of Alignment in Large Language Models
- Multi-step Jailbreaking Privacy Attacks on ChatGPT
- FlipAttack: Jailbreak LLMs via Flipping
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs