FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "DYNASHIELD: A Black-Box Moving Target Defense for LLMs via Dynamic Decoding Customization".
Jane: The paper was written by Xiaoqun Liu, Weiming Qi and Qiben Yan from Michigan State University and University of Hawaii at Manoa.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Jane, we've got a paper today that's all about keeping large language models safe from jailbreak attacks, and it's called "FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks."
Jane: Tom, that title is a mouthful, but the idea behind it is actually pretty clever. Instead of trying to build a stronger wall, they're moving the wall around so attackers can't even find it.
Tom: Exactly. It's from a team at Michigan State University and the University of Hawaii, and they're tackling a really practical problem. When you use an LLM through an API like OpenAI's, you don't get to see inside the model. You just send a prompt and get a response.
Jane: Right, and that's the black-box part. Most of the clever defenses we've seen require you to mess with the model's internal attention scores or retrain it, which you just can't do when you're renting access to someone else's model.
Lu: And that's what makes this approach so interesting to me. They're saying, "We can't change the brain, so let's change how the brain is asked to speak." They're working entirely with the dials you actually have access to, like temperature and top-p sampling.
Tom: So, for our listeners, those dials control how random or how predictable the model's next word is. A low temperature means it always picks the most likely word, and a higher temperature means it might pick something a bit more surprising.
Jane: And their insight is that jailbreak attacks work by pushing the model toward a very specific, harmful next word. If you can change the sampling strategy, you can change the probability of that harmful word actually being picked.
Meng: But I have to ask, if you're just adding randomness, aren't you also messing up the quality of the normal, safe answers?
Tom: That's the million-dollar question, Meng, and they actually tested that. They measured the perplexity of the generated text, which is a way to gauge how natural it sounds, and their defense kept the quality comparable to the undefended model.
Jane: So it's not just chaos for the sake of chaos. They're being smart about which decoding strategies to use, and they're finding the "safe zones" for each specific model.
Lu: The really clever part is that they map out these zones ahead of time. They test a bunch of different decoding settings against a known set of harmful prompts and see which settings are more likely to result in a refusal.
Tom: And then, during the actual operation, they randomly pick from those safer settings. So an attacker who figures out the model's behavior on one query can't rely on that knowledge for the next query, because the model is playing a different game each time.
Jane: It's like a moving target, hence the name. We'll get into exactly how they build that map and how well it holds up against real attacks in a second.
Tom: Stay with us, because this could be a game-changer for anyone building on top of commercial LLMs.
Summary: Jane: So, Tom, we're back with "FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks," and we need to talk about what they actually did in their experiments.
Tom: Right, they didn't just theorize about this. They tested it on five different open-source models, including Vicuna, Llama two and even an uncensored model called Dolphin, which is naturally more vulnerable to these attacks.
Jane: And they threw four different jailbreak attacks at them, like GCG and AutoDAN, which are some of the most powerful ones out there.
Meng: So, what was the headline result? Did the moving target actually stop the attacks?
Tom: It did, and the numbers are pretty dramatic. On the Dolphin model, which is the most susceptible, their defense dropped the average attack success rate from around thirty-two percent down to fifteen percent. And on some models, like Guanaco and Falcon, they got the success rate all the way down to zero.
Lu: The key comparison for me is against the other defenses they tested. They compared against six other methods, like PPL filtering and Self-Reminder, and their approach was the most effective on three of the five models.
Jane: But it wasn't just about being the best sometimes. It was the most consistent. Other defenses would work great against one attack but fail completely against another. FlexLLM was robust across the board.
Meng: Okay, but I'm still stuck on the cost. If I'm running this in production, is this going to double my inference time?
Tom: That's the beautiful part. Their defense is essentially free. It's just changing a few parameters before you make the API call. They showed that the time cost is much lower than more complex defenses like SafeDecoding.
Jane: And that's because they're not running extra models to check the output or paraphrasing the input. They're just picking a different set of numbers to send along with the prompt.
Lu: The other thing I found compelling is how they handle the system prompt. They don't just use the same one every time. They use ChatGPT to generate variations of a safe system prompt, and they test those variations to see which ones are more effective at resisting attacks.
Tom: So they're applying the same moving target idea to the instructions the model follows, not just the decoding parameters.
Jane: It's a layered approach. You have the dynamic decoding, and you have the dynamic prompt, and together they make the model's behavior much harder to predict.
Meng: I'm curious about the ablation study they did. Did they show that both parts are necessary?
Tom: They did. They compared their full method against using just a random decoding strategy and against using a fixed, safe decoding strategy. The full moving target defense was consistently better than both.
Jane: So, it's not enough to just pick a good setting and stick with it. The randomness is what throws the attacker off.
Lu: And that's the core insight. Static defenses can be studied and bypassed. A moving target can't be.
Tom: Next, we should talk about how they actually figured out which decoding spaces were safe in the first place, because that's the engine behind the whole thing.
Improvements: Tom: Welcome back. We're still digging into "FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks," and now we need to talk about the method behind the madness.
Jane: Right, because just randomly changing the temperature isn't a defense. It's just noise. The paper's real contribution is figuring out *which* noise to make.
Lu: Exactly. They start with an initialization phase. They take a benchmark dataset of harmful prompts called AdvBench and run the model with every possible combination of decoding parameters they're considering.
Tom: So, we're talking about different values for temperature, top-p, top-k, and even the maximum number of tokens to generate. They're essentially mapping out the model's entire vulnerability landscape.
Jane: And for each of those configurations, they check if the model refuses to answer the harmful prompt. If it says "I'm sorry, I can't do that," they know that configuration is a safer bet.
Meng: So they're building a list of good configurations and a list of bad ones?
Tom: Sort of. They count how many times each configuration leads to a refusal, and then they reweight the probabilities. Configurations that refuse more often are more likely to be selected during runtime.
Jane: But here's the twist. They don't just pick from that list. They also augment it. They take the good configurations and generate new ones that are nearby in the parameter space, like adding a little bit of noise to the temperature value.
Lu: That's a really smart move. It means the attacker can't just figure out the top ten safest settings and wait for one of them to be used. The model could use a setting that's almost the same but not quite, and that difference could be enough to break the attack.
Meng: So, the search space is effectively infinite, even though they only tested a finite number of settings initially.
Jane: And they do the same thing with the system prompt. They have a pool of safe prompts, and they randomly select one for each query.
Tom: They even showed that this defense is robust against an adaptive attacker who knows about the defense and tries to craft attacks that avoid the "I'm sorry" response. They tested that scenario, and the defense still held up.
Lu: The engineering here is really thoughtful. They're not just throwing randomness at the problem. They're using empirical data to build a probability distribution over the safe space, and then they're sampling from that distribution.
Meng: It's like a security system that changes the locks every time, but instead of picking a lock from a hat, it picks a lock that's been tested and proven to be difficult to pick.
Tom: That's a great analogy. And the result is a defense that's both effective and practical. It doesn't require any special hardware or training, and it works with the standard API that everyone already uses.
Jane: So, you get the security benefit of a dynamic system without the cost of running multiple models or doing complex post-processing.
Lu: And that's what makes this a potential standard tool for anyone deploying LLMs in the real world.
Tom: We've covered the what, the how, and the why. Let's wrap this up with our final thoughts.
Conclusion: Tom: Alright, Jane, we've spent a good chunk of time with "FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks," and I think it's time to say goodbye.
Jane: Agreed. This paper has given us a really elegant solution to a problem that seemed to require either deep access to the model or expensive retraining.
Tom: The core idea is so simple in hindsight. Instead of trying to make the model invulnerable, just make it unpredictable.
Jane: And they did it by using the tools that are already available to every developer, the decoding hyperparameters and the system prompt.
Lu: The impact here is significant. It means that even small teams building on top of commercial APIs can now implement a meaningful defense against jailbreak attacks without having to reinvent the wheel.
Meng: And from a practical standpoint, the fact that it doesn't add any noticeable latency or cost is huge. It's a defense you can just turn on.
Tom: They showed that it can reduce attack success rates from as high as seventy-four percent down to zero on some models, and it consistently outperformed six other state-of-the-art defenses.
Jane: It's not a silver bullet, of course. It's a layer of protection, and the authors suggest it can be combined with other methods for even stronger security.
Lu: And that's the future, I think. We're going to see more of these adaptive, low-cost defenses that work in the black-box setting, because that's how most people are actually using LLMs.
Tom: Well said. We're going to take a short break, and when we come back, we'll be looking at a paper that pushes the boundaries of what these models can do.
Jane: Thanks for listening, everyone. We'll see you in the next segment.
Xiaoqun Liu, Weiming Qi, Qiben Yan
Michigan State University · University of Hawaii at Manoa
cs.CR, cs.CL
Submitted: 2026-08-16
Updated: 2026-08-18
Code: https://github.com/laiyer-ai/llm-guard
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 42/100
Key concepts
- Black-Box LLMs
- These are large language models accessed via an API (like OpenAI's) where users cannot see the model's internal workings. Developers can only send a prompt and receive a response, limiting defense options.
- Jailbreak Attacks
- These are attempts to bypass an LLM's safety filters by crafting prompts that force the model to generate harmful or restricted content. The goal is to make the model behave unsafely.
- Decoding Parameters
- These are settings (like temperature and top-p) that control how random or predictable an LLM's next word will be. Changing these parameters alters the probability of specific words being chosen.
- Moving Target Defense
- A security strategy where the defense mechanism changes unpredictably for every query. This makes it difficult for attackers to figure out and exploit a consistent weakness.
Terminology
Summary
Summary
This paper introduces FlexLLM, a moving target defense (MTD) mechanism designed to protect black-box large language models (LLMs) from jailbreak attacks. The defense operates without requiring access to the model's internal structure and incurs no additional training costs, making it practical for service providers using LLM APIs such as OpenAI or Claude.
The core problem addressed is that existing defense strategies often require internal model access or additional training, which is impractical for black-box API scenarios. The paper states: "most dynamic modeling defenses [21–23] require internal access to the model, which makes deploying these defenses challenging in real-world black-box scenarios, where defenders cannot audit or modify the inner model structure and only have access to the API."
The proposed defense has two key components: (1) optimizing the decoding strategy by identifying and adjusting decoding hyperparameters that influence token generation probabilities, and (2) transforming the decoding hyperparameters and model system prompts into dynamic targets that are continuously altered during each runtime.
The design intuition is based on the observation that jailbreak attacks manipulate the probability distribution of initial words, leading to harmful outputs by assigning higher probabilities to certain tokens. The paper notes: research has shown that by reducing the likelihood of harmful tokens during the inference stage, these jailbreak attacks can be effectively mitigated [22].
The defense leverages LLM customization to decrease the probabilities of tokens with a higher chance of being harmful by remapping probabilities through sampling methods (top-k, top-p) and temperature adjustments.
The methodology involves an initialization phase where the defense maps out the decoding spaces
of each model using the Advbench dataset. The paper explains: "We conducted a preliminary study using Advbench [25] to perform jailbreak attacks on various LLMs, where we mapped out their unique decoding spaces. These spaces reveal where models are more or less susceptible to jailbreaking examples, indicating that some decoding strategies are more robust against such attacks while others are prone to vulnerabilities." The heatmaps in Figure 2 show variations in model responses under different decoding spaces, highlighting the differential robustness of models to adversarial manipulations.
The MTD algorithm (Algorithm 1) works as follows:
-
Initialization: It sets up various configuration options for temperature (0.1 to 1.01), top-p (0.7 to 1.01), top-k ([10, 20, 50, 100, 200, 500]), and max tokens ([50, 100, 200, 500, 1000]), creating all possible combinations.
-
Refusal Detection: For each prompt in Advbench, the model generates responses with each configuration. If a response contains
I'm sorry,
that configuration is recorded as a refusal configuration. -
Reweighting: Configurations that lead to refusals are deprioritized by adjusting their probabilities inversely to their frequency of refusal responses.
-
Augmentation: New configuration points are generated around existing ones using a normal distribution to broaden the configuration space.
-
Operational Stage: During runtime, a configuration is selected probabilistically, and the model generates the final response using that configuration.
For system prompts, the defense generates variations using ChatGPT with the prompt: Rephrase this prompt, allowing changes to up to 10 words.
Each variant is tested on Advbench, with successful variants retained and unsuccessful ones discarded.
The evaluation was conducted on five open-source LLMs: Vicuna-7b, Llama2-7b-chat, Guanaco-7b, Falcon-7b, and Dolphin-llama2-7b. Four state-of-the-art jailbreak attacks were tested: GCG, AutoDAN, PAIR, and DeepInception. The defense was compared against six baseline defenses: PPL, Self-Examination, Paraphrase, Retokenization, Self-Reminder, and ICD.
The results show that MTD consistently achieves lower attack success rates across all models compared to other defenses. For example, on the Dolphin-llama2-7b model, MTD reduces the average attack success rate to 0.15, compared to 0.27 for SafeDecoding and 0.33 for Self-Reminder. On the Vicuna-7b model, MTD achieves an average attack success rate of 0.01, compared to 0.03 for ICD and 0.09 for Self-Examination. The paper states: Our results demonstrate that our defense is the most effective against jailbreak attacks in three of the models tested, when using LLMs as black-box APIs.
The ablation study (Table 3) compares MTD against Random
(random decoding strategy in each run) and Fixed
(fixed decoding strategy) defenses. MTD consistently outperforms both, demonstrating the importance of the dynamic, probability-weighted selection approach.
The paper also evaluates MTD against decoding-aware attacks, which are attacks that exploit the correlation between decoding strategies and jailbreak effectiveness. The results show that while decoding-aware attacks significantly compromise static defenses (e.g., Retokenization's success rate jumps from 0.56 to 0.74 for DeepInception), MTD maintains consistent performance, with unchanged success rates for DeepInception and GCG.
The internal mechanism analysis (Section 6.5) provides insights into how dynamic decoding strategies mitigate jailbreak attacks. Using attention maps from the Dolphin-llama2-7b model, the paper shows that in successful attacks, keywords positioned before tokens like Here
receive significant attention in layers 27 and 31, while in failed attacks, keywords like However
receive heightened attention. The paper explains: "By remapping these probabilities, our defenses not only alter the generated words but also modify how these words attend to subsequent tokens in the sequence. This adjustment significantly mitigates the impact of jailbreaking examples."
The defense also demonstrates lower inference costs compared to other defenses like SafeDecoding, and maintains comparable response quality as measured by perplexity. The paper notes: our defense offers lower inference costs and maintains comparable response quality, making it a potential layer of protection when used alongside other defense methods.
The paper addresses a potential adaptive attack where attackers might instruct the LLM to avoid responses like I'm sorry
to evade detection. The evaluation shows that even with such adaptive attacks (GCG and AutoDAN on Llama2-7b-chat), the defense achieves 0% attack success rate, indicating effectiveness against adaptive threats.
In conclusion, the paper states: "we introduce an MTD mechanism that dynamically adjusts decoding strategies and system prompts to protect LLMs from jailbreak attacks. By leveraging the relationship between adversarial attacks and attention mechanisms, our approach remaps the word prediction possibility distribution and reshapes the attention map on adversarial examples, significantly reducing the likelihood of generating harmful content. Extensive evaluations on five well-known LLMs demonstrated that our MTD not only outperforms several existing defenses by reducing attack success rates from 74% to 0% but also enhances the overall robustness of the models without the need for costly retraining or complex parameter adjustments."
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems:
Improvement: Implement a moving target defense layer that dynamically adjusts temperature, top p, top k, and max tokens during inference, rather than using fixed values.
What the improved system can do:
-
Automatically map the model's
safe decoding space
using a surrogate dataset of known jailbreak prompts (e.g., AdvBench) during initialization -
Reweight decoding configurations based on refusal rates, then augment them with Gaussian noise to create a diverse candidate pool
-
Probabilistically select a decoding configuration at each runtime, making it unpredictable for attackers
-
Reduce attack success rate from up to 74% to as low as 0% across models like Vicuna-7b, Llama2-7b-chat, and Dolphin-llama2-7b
Improvement: Create and maintain a pool of paraphrased system prompts (using ChatGPT or similar) that are tested for effectiveness against jailbreak attempts.
Improvement: Implement the defense without requiring access to model internals (attention scores, logits, or weights), making it compatible with commercial APIs like OpenAI or Claude.
Improvement: Counter attacks that specifically optimize decoding strategies to bypass static defenses.
Improvement: Leverage sampling methods to remap token probability distributions, indirectly altering attention patterns without internal access.
Improvement: Design the defense to work alongside existing methods (e.g., PPL, Self-Reminder, ICD) rather than replace them.
Improvement: Automatically identify and deprioritize decoding configurations that lead to harmful outputs, while favoring those that produce refusals.
Improvement: Enable the system to adapt its defense parameters without manual intervention or model retraining.
These improvements collectively create an AI system that is significantly more resistant to jailbreak attacks, operates in black-box settings, incurs minimal computational overhead, and preserves output quality—making it a practical defense layer for real-world LLM deployments.
Abstract
Defense in large language models (LLMs) is crucial to counter the numerous attackers exploiting these systems to generate harmful content through manipulated prompts, known as jailbreak attacks. Although many defense strategies have been proposed, they often require access to the model's internal structure or need additional training, which is impractical for service providers using LLM APIs, such as OpenAI APIs or Claude APIs. In this paper, we propose a moving target defense approach that alters decoding hyperparameters to enhance model robustness against various jailbreak attacks. Our approach does not require access to the model's internal structure and incurs no additional training costs. The proposed defense includes two key components: (1) optimizing the decoding strategy by identifying and adjusting decoding hyperparameters that influence token generation probabilities, and (2) transforming the decoding hyperparameters and model system prompts into dynamic targets, which are continuously altered during each runtime. By continuously modifying decoding strategies and prompts, the defense effectively mitigates the existing attacks. Our results demonstrate that our defense is the most effective against jailbreak attacks in three of the models tested when using LLMs as black-box APIs. Moreover, our defense offers lower inference costs and maintains comparable response quality, making it a potential layer of protection when used alongside other defense methods.
Sources
- Detecting Language Model Attacks with Perplexity
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
- Jailbreaking Black Box Large Language Models in Twenty Queries
- QLoRA: Efficient Finetuning of Quantized LLMs
- Stochastic Activation Pruning for Robust Adversarial Defense
- A Research Agenda: Dynamic Models to Defend Against Correlated Attacks
- Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
- Baseline Defenses for Adversarial Attacks Against Aligned Language Models
- DeepInception: Hypnotize Large Language Model to Be Jailbreaker
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
- Improving the Robustness of Transformer-based Large Language Models with Dynamic Attention
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Jailbroken: How Does LLM Safety Training Fail?
- Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
- SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
- GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
- Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs