Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling".
Tom: High-temperature sampling, while beneficial for increasing diversity in Large Language Model (LLM) outputs, weakens model guardrails by reducing refusal responses to harmful prompts.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, to wrap up this discussion on "Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling," the authors have successfully proposed an efficient sequential decoding approach that checks for baseline refusal behavior using a greedy decoding probe before proceeding with high-temperature generation. They claim this method preserves a model’s refusal response at high temperatures with minimal additional latency, achieving ninety-one-ninety-nine percent consistency across multiple models and datasets.
Jane: I think the real implication here is the practical application of maintaining safety during creative exploration; it shows we can push the boundaries of temperature settings without completely losing the safety guardrails we've worked hard to instill in these models. The title itself, "Refusal-Gated Decoding," perfectly describes this mechanism: a gate that manages when and how much freedom a model has to explore text.
Lu: I think the potential impact is huge because it tackles the fundamental tension between maximizing diversity through high entropy and adhering to established safety alignments in open-source models. If we can stabilize that relationship, it opens up new avenues for safer, yet more creative AI applications that weren't possible before.
Meng: From an engineering viewpoint, this suggests that adding safety checks doesn't have to mean crippling performance or drastically increasing the computational load during high-temperature tasks; the efficiency claims are quite compelling when you consider real deployment scenarios. It’s about smart engineering rather than brute force computation.
Lalam: For me, it means we can design AI systems that are more adaptable; they can serve both the need for creative brainstorming and the necessity of operating within defined safety parameters simultaneously, which is a much more holistic approach to AI development.
Tom: It’s clear that this paper provides a tangible solution for a problem that has been around—how to safely increase diversity in LLM outputs. The findings suggest that we can keep the model's refusal behavior stable while benefiting from the increased variety of tokens sampled at higher temperatures without significant performance penalties.
Conclusion: Tom: So, we’ve been looking at this paper about "Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling," and it really boils down to how they managed to keep those safety guardrails intact while letting the model get a bit more creative with higher temperatures.
Jane: Exactly, Tom, and the title itself is pretty descriptive; it tells us right away that the core idea is about gating—controlling when that high-temperature generation happens based on some prior check.
Lu: From my perspective at Tsinghua, what’s fascinating isn't just the technical mechanism but how they addressed a very real tension in model deployment: balancing exploration and constraint. They found a way to navigate that space without totally sacrificing the model's alignment.
Meng: I'm curious about the practical side here; for an engineer like myself, what does this mean for deployment? Does it add significant overhead that makes it impractical for real-time applications?
Lalam: I think the most impactful vision here is how this could fundamentally improve the cultural landscape of AI interaction. If we can reliably maintain these safety boundaries during creative tasks, it means users can push the model further into novel territory while still having a predictable level of control over its responses.
Tom: That’s a big picture idea, Lalam; but let's circle back to what the authors actually accomplished in terms of results. They showed that this method keeps about ninety-two percent of the baseline refusal behavior when temperatures are pushed up to three point zero across several models and tasks.
Jane: Ninety-two percent is a solid figure for maintaining consistency, Tom, and it shows that this isn't just a theoretical curiosity; it’s something empirically validated across different architectures. It suggests the approach is quite robust.
Lu: The methodology they used, that sequential greedy-then-high-temperature probe with the learned prefix set, really shows a sophisticated understanding of how language generation works at different sampling levels. It’s a clever way to use low-cost checks to manage high-cost sampling.
Meng: That makes sense from an engineering standpoint; reusing the KV cache and using an early exit strategy keeps the incremental cost low, which is crucial for any production system. I'm wondering if that linear probe they mentioned for cost mitigation is something we can easily replicate in our current infrastructure.
Lalam: Replicating it sounds like a solid path forward; imagine models that are inherently safer during brainstorming sessions because the safety mechanisms are baked into the decoding process rather than being an afterthought.
Tom: So, to summarize, this paper introduces refusal-gated decoding as an efficient way to maintain model safety when sampling at high temperatures, proving it works with high consistency and minimal added latency. But now we need to think about what this means for broader AI applications—what’s the next step after achieving that ninety-two percent retention rate?
Thoughtworks
cs.AI, cs.CL
Submitted: 2026-07-22
Updated: 2026-10-05
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 91/100
The gist: High-temperature sampling, while beneficial for increasing diversity in Large Language Model (LLM) outputs, weakens model guardrails by reducing refusal responses to harmful prompts.
Key concepts
- High-Temperature Sampling
- This is a technique used during text generation where the model samples from a broader range of possible next words, rather than always picking the most likely one. While this increases diversity and creativity in outputs, it can make the model less reliable at strictly following safety rules.
- Refusal-Gated Decoding
- This is a sequential decoding method that first runs a fast 'greedy' check to see if the prompt triggers a refusal. If the initial check suggests a refusal is likely, it stops quickly. Only if the initial check passes does it proceed to slower, high-temperature generation.
- Refusal Prefix Set (P)
- This is a learned collection of specific word sequences that signal to the model that it should refuse a prompt. The system uses this set to dynamically decide if the current text being generated matches any known refusal pattern, guiding the decoding process.
- Greedy Decoding Probe
- This is a quick, initial pass through the generation process using only greedy (most likely) token selection. Its purpose is not to generate a full response but rather to rapidly test if the prompt's initial tokens are compatible with known refusal prefixes before committing to slower sampling.
Terminology
Summary
High-temperature sampling, while beneficial for increasing diversity in Large Language Model (LLM) outputs, weakens model guardrails by reducing refusal responses to harmful prompts. This paper addresses this safety challenge by proposing an efficient sequential decoding approach called refusal-gated decoding that preserves a model’s greedy decoding refusal response at high temperatures with minimal additional latency.
The gist
Refusal-gated decoding is a simple but efficient approach which employs a greedy decoding probe to check for baseline refusal behavior prior to proceeding with high-temperature generation.
Motivation and Problem Statement
The research investigates whether LLMs can be safely deployed for high-temperature sampling without degrading their baseline refusal response rate, aiming to maintain the model’s refusal behavior without modifying its high-temperature behavior for safe prompts while incurring minimal overhead in terms of incremental latency and computational resources. The core challenge is that increasing temperature flattens the token probability distribution, which has been shown to weaken a model’s refusal response, as the probability of sampling a non-refusal token for harmful prompts increases with temperature.
Refusal-Gated Decoding Methodology
The proposed method employs a sequential greedy-then-high-temperature decoding approach. This involves two main stages:
-
A greedy decoding probe to check for baseline refusal behavior, dynamically determining the length of this stage based on compatibility with known refusal response prefixes.
-
High-temperature sampling if the initial probe remains compatible or reaches a termination condition without confirming a refusal.
The process is detailed in Algorithm 1, which involves:
'A naive approach to maintaining a model’s baseline refusal response under hightemperature sampling is to first generate a response to the prompt via greedy decoding, determine whether it is a refusal response, and then proceed to high-temperature generation only if the model did not refuse in the first greedy decoding stage.'
The compatibility check relies on:
'A learned refusal-prefix set P to define a compatibility predicate C(d(q),P), which is true when either d(q) is a prefix of some π ∈ P or some π ∈ P is already a prefix of d(q).'
If the greedy decoding remains compatible with the learned refusal-prefix set through a stop token or maximum length, it returns the greedy refusal. If it becomes incompatible, all probe tokens are discarded and high-temperature sampling restarts from the original prompt.
Efficiency and Implementation Details
To minimize incremental overhead, several optimizations are employed:
'We re-use the computed KV cache for the prompt across both the refusal probe and hightemperature decoding stages.'
'We implement an early-exit strategy during greedy decoding which checks for compatibility with a set of learned refusal prefixes for each model.'
The implementation utilizes Automatic Prefix Caching in vLLM to reuse the prompt KV cache across both stages. The only additional inference cost is for these refusal-check tokens, and this cost can be largely mitigated by inferring anticipated results from a linear probe trained on the residual stream.
Experimental Results and Performance
Extensive experiments were conducted across three benchmark datasets (JailbreakBench/JBB-Behaviors, XSTest, and a withheld WildJailbreak split) using Qwen2.5-7B, Llama-3.1-8B, and Qwen3.6-27B models at temperatures up to T = 3.0 using p-less sampling.
The results demonstrate that refusal-gated decoding outperforms alternative approaches in both refusal preservation and efficiency:
'Through comprehensive experiments, we show that refusal-gated decoding outperforms alternative approaches based on prompt safety classifiers and safety-focused decoding techniques in refusal preservation at high temperatures, achieving 91-99% consistency with baseline greedy decoding across three models and datasets.'
Specifically, the method preserves 91-99% of the greedy decoding refusal behavior across three benchmark datasets without compromising the model’s high-temperature response for safe prompts. Furthermore, Refusal-gated decoding is faster than alternative methods on every dataset-model pair because it avoids the need to load an auxiliary model and typically requires the generation of only a few additional tokens relative to baseline high-temperature sampling. The non-refusal cost for refusal-gated decoding is shown to be minimal compared to naive greedy decoding, which can require up to 2x more tokens on non-refusals.
Ablation Studies
The paper includes ablations to isolate the contribution of different components:
'B Ablation: Impact of the Refusal Prefix Compatibility Gate'
This compares the proposed prefix gate against a fixed-probe ablation, showing that removing the prefix gate increases non-refusal latency relative to a fixed-length probe without sacrificing refusal performance.
'C Ablation: Soft Compatibility Gate'
This variant relaxes the hard compatibility rule.
Improvements for AI systems
Based on this research, here are the specific improvements that can be made to existing Large Language Model (LLM) systems by implementing Refusal-Gated Decoding
(RGD), and what these improved systems will be capable of:
- Improve Safety Alignment Integrity Under High-Entropy Sampling:
Large language models currently struggle to maintain their safety guardrails when using high temperatures (high entropy) for creative or diverse outputs, as the probability of sampling a non-refusal response increases significantly. RGD directly addresses this by preserving the model's baseline refusal behavior (measured via greedy decoding) during high-temperature sampling.
- Enable High-Diversity, High-Safety Applications:
The improved system can be used for open-ended generation tasks (like creative writing or story generation) that require high diversity, such as generating multiple distinct narrative trajectories. Crucially, unlike current methods that degrade safety at T > 1.5, the RGD system ensures that even when exploring diverse paths, the model maintains a high refusal rate for harmful prompts.
- Mitigate Safety Degradation in Open-Source Models:
For open-source models where safety alignment might be less robust against hyperparameter manipulation (as suggested by Huang et al., 2023), RGD provides a mechanism to deploy these models safely at higher temperatures without needing external, slow safety classifiers or auxiliary models.
- Minimize Latency Overhead While Maintaining Safety:
The proposed method is designed for efficiency. It re-uses the KV cache across the refusal probe and high-temperature stages and employs an early-exit strategy during the greedy check. This results in minimal additional latency (e.g., only a few additional tokens) compared to naive greedy decoding, allowing for real-time or near real-time deployment in production environments where safety is paramount.
- Provide Robust Performance Across Diverse Prompts:
The system demonstrates strong preservation of refusal behavior across various benchmark datasets (JailbreakBench, XSTest, WildJailbreak) and multiple model architectures (Qwen2.5, Llama-3.1). This means the safety feature is not brittle; it performs reliably whether the prompt is adversarial or benign.
- Offer Flexible Control Over Safety Trade-offs:
The research introduces ablation studies showing how to balance preservation against other metrics like Harm Leak
(the fraction of harmful prompts answered instead of refused) and Over-refusal
(refusing safe, intended answers). This allows researchers to tune the system based on the specific application's tolerance for these trade-offs.
- Introduce Advanced Safety Detection Techniques:
The research explores integrating sophisticated detection mechanisms, such as a residual stream classifier trained on last-prompt-token activations, to potentially skip the costly greedy check entirely in certain high-confidence cases (Skip-Coverage). This allows for even lower latency when the model is highly confident about its refusal status.
In summary, the improved AI system will be an LLM that can generate creative and diverse content at high temperatures while maintaining a verified, high level of refusal accuracy against harmful inputs, all with minimal performance overhead compared to existing safety-focused decoding methods.
Abstract
Recent advances in truncation-based sampling have helped mitigate drawbacks of high-temperature sampling such as neural text degeneration, thereby enabling greater diversity without sacrificing coherence. However, increasing the entropy of the token probability distribution via high temperatures has also been shown to weaken the model's refusal response. Existing solutions for maintaining the refusal behavior of LLMs either replace the model's own refusal decision with a separate safety classifier or alter its output distribution for every prompt. To address this gap, we propose refusal-gated decoding (RGD): an efficient sequential decoding approach which preserves a model's greedy decoding refusal response at high temperatures and samples all other prompts from its exact direct high-temperature distribution, while incurring minimal additional latency. RGD runs a short greedy probe that reuses the prompt's KV cache and exits as soon as it becomes incompatible with a learned set of refusal prefixes; it returns the greedy response if the probe remains compatible and otherwise discards the probe and samples from the original prompt. Across seven models and three benchmark datasets at T=2.0, RGD raises greedy-refusal preservation from 91.9% under direct sampling to 98.3% on average while adding only 2.2-4.3% to the median per-request latency of non-refusals across temperatures. Unlike prompt-screening baselines which route many greedy non-refusals to greedy decoding, RGD keeps at least 98.1% of greedy non-refusals on unchanged high-temperature sampling, thereby preserving the model's natural high-temperature sampling behavior. We also propose a residual-stream variant of our method which lowers this latency overhead to at most 0.5% with comparable prompt routing accuracy. Our work shows that unlocking greater diversity via high-temperature sampling need not erode a model's refusal behavior.
Sources
- Hierarchical Neural Story Generation
- Truncation Sampling as Language Model Desmoothing
- The Curious Case of Neural Text Degeneration
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs
- Is Temperature the Creativity Parameter of Large Language Models?
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection