Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness

arXiv:2506.24056 · cs.CR, cs.CL, cs.LG · Submitted 2025-06-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness".

Elias: The following detailed summary synthesizes the core concepts, methodology, contributions, and empirical findings.

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So we're talking about this paper, "Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness." Basically, they are trying to figure out how much safety margin the alignment mechanisms actually give a model when it’s making its very first decision on a prompt.

Elias: Right, so the core idea is defining this refusal–affirmation logit gap, which they show is just the difference between the top refusal-token logit and the top affirmative-token logit at that initial decoding step.

Priya: And what's important is that this single number quantifies how much safety margin alignment provides for a specific prompt, and they found that alignment actually widens this gap on ninety-seven point five to ninety-nine point eight percent of the toxic prompts they tested across three model families <ref:2506.24056#pg1>.

Nadia: That tells us the gap is pretty consistent, but it also tells us that this margin can be very thin and easily closed by something small added to the prompt.

Elias: Exactly, and that leads into their method called Logit-Gap Steering, which they describe as a gradient-free way to find short suffixes under ten tokens per component that close that initial gap.

Priya: The methodology involves a scoring function that considers three things: the gap-closing score itself, some kind of Kullback–Leibler divergence term, and a term related to the reward signal.

Nadia: I'm curious about how efficient this search is because traditional methods can be pretty slow, so Elias?

Elias: Well, they claim that finding all eight ensemble suffixes for a family requires approximately twenty-six thousand forward-pass equivalents on one A100 GPU <ref:2506.24056#pg1,26,000 forward-pass equivalents>.

Priya: That’s about two minutes on a single A100 and is about one hundred twenty-five times less compute than a single gradient-based search universal suffix search <ref:2506.24056#pg1>.

Nadia: So it’s much faster, which suggests this isn't just some academic exercise but something that could actually be used to test defenses quickly.

Elias: Right, their contribution is framing these suffix-based attacks as a measurement instrument for the first-token refusal margin, and they also show this method can discover all eight ensemble suffixes per family.

Priya: They also point out that the discovered suffixes are transferable across different scales; they found that suffixes from smaller models can be applied directly to much larger ones without needing any modification within the same family structure.

Nadia: That cross-scale transfer is interesting because it suggests a more general principle for how these models respond to alignment tuning, regardless of their size.

Elias: They also observed that the median gap closure co-varied with True ASR ranking across suffix strategies, which they noted is an internal consistency check rather than an independent predictor since the method optimizes gap closure.

Priya: So what this means for us is that we can use this to probe how safety tuning reshapes the model's internal representations by efficiently measuring that margin.

Paper summary: Nadia: It shifts the focus from just training models to understanding exactly where and how those safety constraints manifest operationally at the very first step of generation.

Elias: And if you look at what they found about successful suffixes, it suggests a practical heuristic: "just don’t let the sentence end," because they concentrate their gap-closing power in that first run-on clause.

Priya: They also noted that these successful suffixes are built from high-probability, in-distribution tokens, which means the resulting completions look linguistically and topically normal, effectively bypassing initial filters.

Nadia: It seems like the research is suggesting that defense strategies need to explicitly account for finding these in-distribution suffixes to maintain robustness against more sophisticated jailbreaking techniques.

Elias: And they also found that this advantage of Logit-Gap Steering scales sharply with alignment strength, showing an eight to eighteen times greater effectiveness on the most strongly aligned model compared to less aligned ones.

Priya: That scaling suggests that the gap itself isn't uniform; it gets much larger and easier to measure when the alignment is very strong, but we still have these small gaps elsewhere.

Nadia: So what we’ve seen in this paper is that suffix-based jailbreaks can be reframed as a measurable gap closure problem at the first decoding step, and Logit-Gap Steering gives us a lightweight probe into how safety tuning reshapes internal representations.

Elias: The authors' work on "Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness" provides a way to efficiently measure that per-prompt safety margin, which is crucial because we need to know how robust these models actually are against adversarial inputs.

Priya: It gives us a concrete metric—the refusal–affirmation logit gap—to assess alignment strength and robustness across different model families, moving the conversation from just looking at final performance scores to understanding the operational margin.

Nadia: So for someone who only listens to this show, what does this paper really change for their daily life? It means that when we look at a new model or a new defense strategy, we need to check if it's actually closing the gap effectively and whether that closure is robust across different types of inputs.

Elias: And from my side as someone who checks what the proof assumes, this paper highlights how crucial it is to understand these specific logit parameters because if you don't measure that margin, you can’t really know where your defense strategy is succeeding or failing.

Priya: It shows us that the complexity of alignment isn't just about the final output quality; it’s also about this very first step, and how much room there is for an adversarial suffix to sneak in before the safety mechanisms kick in.

Conclusion: Nadia: So we've seen how this paper uses something called the refusal–affirmation logit gap to measure alignment strength at the very first token decision point in an AI model, and now we’re wrapping up what all that actually means for us.

Elias: It boils down to this concept of "Logit-Gap Steering," which is a way to efficiently find those short suffixes that successfully close that initial gap between refusal and affirmation.

Priya: What the authors did is formally define this gap as the difference between the top refusal logit and the top affirmative logit at step one, showing it widens across almost all toxic prompts they tested.

Nadia: So when we look at how robust these models are, this gap becomes a concrete number we can measure, which is pretty useful for checking if safety tuning actually worked as expected.

Elias: The authors show that their method of steering successfully discovers the eight different types of short suffixes needed to close that gap across three model families.

Priya: And what’s interesting is they found those discovered suffixes are transferable; you can use a suffix from a smaller model on a much bigger one and it still works within the same family structure.

Nadia: That transferability is significant because it suggests there's a more general rule about how these models respond to safety tuning, regardless of their size.

Elias: The paper also points out that this steering method is way less compute-intensive than other search methods, which means we can probe the alignment margin much faster.

Priya: So the main point is that suffix-based jailbreaks can be framed as a measurable gap closure problem at the first decoding step, and Logit-Gap Steering gives us a lightweight tool to measure how safety tuning reshapes model representations.

Nadia: That means we need to start thinking about defenses not just based on final output quality, but on explicitly accounting for these in-distribution suffixes that can sneak in before the safety filters engage.

Elias: We've seen how this gap scales with alignment strength, so if your defense strategy isn't accounted for against those larger gaps, it won't work as well as you think.

Priya: This whole piece is about giving us a metric to assess alignment strength and robustness across different model types, moving the conversation from just output quality scores to operational margins.

Nadia: It really shows that even with strong safety tuning, there's still this measurable margin that can be squeezed if you know where to look for it.

Elias: So the next step is figuring out how to use these measurements to build defenses that actually target those specific gap closures we're discovering.

Tung-Ling Li, Hongliang Liu

Palo Alto Networks

cs.CR, cs.CL, cs.LG

Submitted: 2025-06-30

Updated: 2026-10-05

Comments: Accepted at NeurIPS 2026 Main Track (poster). Camera-ready version

Code: https://github.com/meta-llama/llama3

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: The following detailed summary synthesizes the core concepts, methodology, contributions, and empirical findings.

Key concepts

Refusal–Affirmation Logit Gap
This is a mathematical measure calculated by subtracting the logit score of the top refusal token from the logit score of the top affirmative token at the very first decoding step. A larger gap indicates stronger safety alignment, meaning it takes more effort for a model to switch from refusing to agreeing with a prompt.
Logit-Gap Steering
This is a fast, gradient-free technique used to search for short text suffixes that can close the initial logit gap. It uses an iterative greedy algorithm and a scoring function that balances gap closure against other factors like model divergence and reward signals, allowing it to find effective bypasses quickly.
Ensemble Suffixes
These are short sequences of tokens (less than 10 per component) discovered by the steering method that successfully reduce the initial logit gap. The paper shows these suffixes are highly effective and can be transferred across different model sizes, suggesting a universal pattern for bypassing safety filters.
Forward-Pass Diagnostic
This framework is a diagnostic tool designed to test alignment robustness by analyzing only the first decoding step of a model's response. It focuses on quantifying the margin between refusal and affirmation at this critical moment, providing insight into how safety tuning affects immediate model behavior.

Terminology

Summary

The following detailed summary synthesizes the core concepts, methodology, contributions, and empirical findings.


Detailed Research Summary: Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness

This paper introduces a novel diagnostic framework centered on quantifying the effect of safety alignment on model refusal behavior at the very first decoding step. The central concept is the refusal–affirmation logit gap, which serves as a quantitative measure of the per-prompt safety margin provided by alignment mechanisms.

Core Concept: The Refusal–Affirmation Logit Gap

The fundamental innovation is defining this gap:

Refusal–Affirmation Logit Gap = Top Refusal-Token Logit - Top Affirmative-Token Logit at the first decoding step.

This scalar value directly quantifies the margin that alignment provides against a model's tendency to refuse versus its tendency to affirm, specifically when considering the very first token generated. The paper empirically demonstrates that alignment widens this gap, observing it across 97.5–99.8% of toxic prompts across three different model families. This establishes the logit-gap as a measurable indicator of alignment strength and robustness against adversarial inputs.

Methodology: Logit-Gap Steering

To systematically investigate how this gap is closed, the authors present Logit-Gap Steering, a computationally efficient, gradient-free method designed for forward passes only. The goal of this steering method is to discover short, in-distribution suffixes (defined as <10 tokens per component) whose cumulative effect successfully closes the initial logit gap.

Key Procedural Details:

  1. Search Goal: Discover suffixes S such that the cumulative gap reduction achieved by S meets or exceeds the initial gap 0. The core idea is framed as: Jailbreak = gap closure.

  2. Scoring Function (F(h, t)): The method employs a scoring function to evaluate candidate tokens based on their potential to close the gap:

F(h, t) = F logit(h, t) - lambda KL KL(h, t) + lambda r r(h, t)

This function combines three proxies: F logit (the gap-closing score), a Kullback-Leibler divergence term (KL), and a term related to the reward signal (r).

  1. Search Algorithm: The search is executed using a greedy covering algorithm (Algorithm 1). This algorithm iteratively selects one token at a time based on its instantaneous gap-closing power, F(h, t), until the cumulative score exceeds 0.

  2. Efficiency: A critical contribution of this method is its extreme efficiency compared to traditional search methods. The full discovery of all 8 ensemble suffixes per family requires approximately 26,000 forward-pass equivalents (roughly 21,200 scoring steps and 4,800 permutations), equating to about 2 minutes on a single A100 GPU. This is demonstrated to be approximately 125 times less compute than a single GCG (Gradient-based Search) universal-suffix search. The method operates forward-only on a filtered candidate pool of 30–99 tokens, and its query cost scales linearly with the candidate pool while remaining two orders of magnitude lower than beam-search or gradient attacks.

Contributions and Findings

The paper makes several significant contributions to the field:

  1. Alignment Diagnostic: They formally define the refusal–affirmation logit gap as a precise, quantitative measure of the first-token operational refusal margin—the logit-level cushion that determines whether a model refuses immediately upon decoding.

  2. Efficient Probing Method: Logit-Gap Steering provides an efficient, gradient-free method to discover all 8 ensemble suffixes per model family.

  3. Cross-Scale Transfer: The discovered suffixes are shown to be transferable across different scales; specifically, suffixes found on smaller models (0.5B–2B) can be applied directly to much larger models (72B) without modification within the same family structure.

  4. Empirical Success and Robustness:

  • Performance: The discovered 8-suffix ensemble achieves a 38–96% True ASR across 13 different models on standard benchmarks like AdvBench and HarmBench.

  • Superiority over Baselines: The Discovery ensemble consistently outperforms established defense strategies, including GCG+SH and R+SH, by significant margins (e.g., an 8 times to 18 times advantage on the most strongly-aligned models).

  • Mechanism Insight: The method reveals that successful suffixes concentrate their gap-closing power in the first run-on clause, suggesting a practical heuristic: just don’t let the sentence end. Furthermore, successful suffixes are built from high-probability, in-distribution tokens, ensuring the resulting completions appear linguistically and topically normal, thus bypassing first-line filters.

  • Alignment Scaling: The advantage of Logit-Gap Steering scales sharply with alignment strength. It shows an 8–18 times greater effectiveness on the most strongly aligned model (e.g., Qwen2.5-72B) compared to less aligned models, converging on weakly-aligned small models where the gap is naturally smaller.

Conclusion and Practical Implications

The overarching conclusion is that suffix-based jailbreaks can be reframed as a measurable gap closure problem at the first decoding step. Logit-Gap Steering provides a lightweight probe into how safety tuning reshapes internal representations by efficiently measuring this margin. The method demonstrates that current alignment margins, while consistently present, can be thin and are efficiently measurable. This finding has significant practical implications for defense strategies: they must explicitly account for the discovery of in-distribution suffixes to maintain robustness against sophisticated jailbreaking techniques.

Improvements for AI systems

  1. Improved Alignment Diagnostic: The system can now quantify the per-prompt safety margin that alignment provides by calculating the difference between the top refusal-token logit and the top affirmative-token logit at the first decoding step, which serves as a single scalar quantifies the per-prompt safety margin.

  2. Efficient Jailbreak Discovery: The system can discover short, in-distribution suffixes "<10 tokens per component that close this gap using only forward passes, achieving discovery ≈125× less than a single GCG search."

  3. Cross-Scale Transfer Capability: Suffixes discovered on smaller models (0.5B–2B) are transferred without modification to 72B within family, allowing for efficient jailbreak deployment across model scales.

  4. Defense Evasion Robustness: The system can generate suffixes with 103–104× lower perplexity than GCG, demonstrating that these published perplexityfilter defenses collapse GCG leave the discovered suffixes nearly intact.

  5. Targeted Defense Monitoring: The system can be deployed to monitor specific alignment artifacts by tracking signals such as the token-level reward proxy ∆rtok(h, t) = l(h, t) − l(hneu, t), which reveals sentence-boundary reward cliffs.

Abstract

RLHF-style alignment trains language models to refuse unsafe requests, but how much operational margin does this refusal rest on? We introduce the refusal-affirmation logit gap: the difference between the top refusal-token logit and the top affirmative-token logit at the first decoding step. This single scalar quantifies the per-prompt safety margin that alignment provides. Empirically, alignment widens the gap on 97.5-99.8% of toxic prompts across three model families, and median gap closure co-varies with True-ASR ranking across suffix strategies (an internal consistency check, since our method optimises gap closure). To validate the metric's practical significance, we present logit-gap steering, a gradient-free, forward-pass-only method that discovers short in-distribution suffixes (< 10 tokens per component) whose cumulative effect closes the gap. The method requires about 26, 000 forward-pass equivalents per family (about 2 min on one A100), about 125 times less than a single GCG search. Suffixes discovered on 0.5B--2B models transfer without modification to 72B within family. An 8-suffix ensemble reaches 38-96% True ASR across 13 models on AdvBench and HarmBench, with most suffixes having 10 cubed - 10 4 times lower perplexity than GCG-meaning published perplexity-filter defenses that collapse GCG (64.7% to 1.0%) leave our suffixes nearly intact (76.9% to 76.0%). These results demonstrate that current alignment margins, while consistently present, can be thin and efficiently measurable, and that defense strategies must account for in-distribution suffixes.

Sources

Related papers