Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations

arXiv:2605.28553 · cs.AI, cs.CR · Submitted 2026-05-27 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations".

Jane: The paper was written by Matteo Gioele Collu, Riccardo Conte, Alberto Giaretta, Denis Kleyko, Mauro Conti et al. from University of Padua, Italy and Örebro University, Sweden and Fondazione Bruno Kessler, Italy.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary of Findings: Tom: We've covered so much ground, from how refusal signals exist in the core of "Refusal Before Decoding" to the impressive efficiency gains of Mechanistic AutoDAN.

Jane: It’s truly a comprehensive look at AI behavior, showing us that these models are more structured and less random than we might have thought about their internal workings.

Lu: The exploration into how these signals transfer across different architectures is crucial for understanding the generalized capabilities of modern AI systems.

Meng: I think the most important practical finding is that because we can predict refusal before decoding, we can build highly targeted and efficient monitoring tools to ensure safety in deployment.

Lalam: I hope that by identifying these vulnerabilities, this work helps us design a more trustworthy relationship with AI for all of us.

Tom: Before we wrap up, Lu, any final thoughts on the theoretical implications?

Lu: I think it's proof that the architecture itself holds the potential for understanding complex behavior and how it guides our use of AI.

Meng: Just to add to that, I hope this provides us with concrete tools for better defense strategies in a real-world setting where we need them.

Lalam: For me, it's about building a culture of understanding and responsibility around AI development as these models get more powerful.

Mechanistic AutoDAN and Improvements: Tom: We’ve established *that* refusal signals exist early on, but now the authors show us *how* to use them in "Refusal Before Decoding" with this new method called Mechanistic AutoDAN.

Jane: It’s a clever modification of the standard AutoDAN technique, which is a genetic algorithm for jailbreaking prompts. The authors replaced the traditional method of evaluating success by checking if the final output matches a target string.

Meng: That’s the practical improvement I'm looking at; instead of running the entire model through all layers just to see if it gives us what we want, we only run it up to a certain point, using those intermediate activations as our score.

Lu: This suggests that by focusing on the internal representation, we are making the search problem much more efficient because we aren't wasting time waiting for the final token generation.

Lalam: Efficiency is crucial for us; if AI can be optimized to find its weaknesses faster, it becomes a tool that we can address and refine more effectively in cultural use.

Tom: And Jane, when comparing Mechanistic AutoDAN to the original AutoDAN, the authors report significant gains in speed—up to seventy-two percent reduction in search time per iteration.

Meng: That seventy-two percent is massive for an iterative process like a genetic algorithm; it drastically changes how long an attack takes and makes it more practical.

Lu: It makes sense that this optimization is possible because the internal signal already contains the information we need, allowing us to skip the later, computationally expensive stages of evaluation.

Implications and Usefulness: Tom: Now, moving from speed, let's talk about usefulness. The authors found that Mechanistic AutoDAN isn't a universal fix; its effectiveness depends on several factors.

Jane: They noted that the usefulness of the probe-guided search actually increases as the model gets larger and more robust in general terms.

Lu: That’s a very interesting theoretical implication, Jane—that safety behaviors might be more clearly defined or structured in larger models, providing clearer signals for optimization than smaller ones.

Meng: From an engineering viewpoint, that tells us where to focus our resources; if we want the best results from an attack or defense against it, we should prioritize the larger architectures where the signal is strongest.

Lalam: It suggests a hierarchy of safety and vulnerability in AI, which is something I think society needs to be mindful of—that not all's security challenges are equal.

Tom: And Jane, when looking at the results across different models, they found that certain blocks in specific models provide much better guidance than others.

Jane: It seems like the timing matters; a signal might be present early on, but it only becomes actionable for our search if we wait until enough layers have processed it and developed the information.

Lu: The fact that the strongest transferability results come from intermediate activations also suggests that those layers are capturing concepts that are less tied to specific model implementations.

Conclusion and Final Thoughts: Tom: We've covered so much ground, from how refusal signals exist in the core of "Refusal Before Decoding" to the impressive efficiency gains of Mechanistic AutoDAN.

Jane: It’s truly a comprehensive look at AI behavior, showing us that these models are more structured and less random than we might have thought about their internal workings.

Lu: The exploration into how these signals transfer across different architectures is crucial for understanding the generalized capabilities of modern AI systems.

Meng: I think the most important practical finding is that because we can predict refusal before decoding, we can build highly targeted and efficient monitoring tools to ensure safety in deployment.

Lalam: I hope that by identifying these vulnerabilities, this work helps us design a more trustworthy relationship with AI for all of us.

Tom: Before we wrap up, Lu, any final thoughts on the theoretical implications?

Lu: I think it's proof that the architecture itself holds the potential for understanding complex behavior and how it guides our use of AI.

Meng: Just to add to that, I hope this provides us with concrete tools for better defense strategies in a real-world setting where we need them.

Lalam: For me, it's about building a culture of understanding and responsibility around AI development as these models get more powerful.

Tom: Thank you all so much for breaking down this research with us today! We’re wrapping up our discussion on "Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations."

Jane: A huge thank you to Lu, Meng, and Lalam as well.

University of Padua, Italy · Örebro University, Sweden · Fondazione Bruno Kessler, Italy

cs.AI, cs.CR

Submitted: 2026-05-27

Updated: 2026-09-03

Code: https://github.com/tatsu-lab/stanford_alpaca

Importance score: 87/100

The gist: As a diligent researcher operating under high stakes, I must inform you that while you have provided the title of the paper—"Refusal Before Decoding: Detecting and Exploiting Refusal Signals in

Key concepts

Refusal Signals
These are signals present in the core of an LLM indicating its refusal to perform a task. The research focuses on detecting these signals early in the model's processing stages before the final output is generated.
Mechanistic AutoDAN
This is a modification of the standard AutoDAN technique, which is a genetic algorithm for jailbreaking prompts. It replaces checking if the final output matches a target string by using intermediate activations as a score.
Intermediate Activations
These are internal representations within an LLM. Utilizing these allows researchers to predict refusal before decoding, making the search process much more efficient because they skip computationally expensive later stages.

Terminology

Summary

As a diligent researcher operating under high stakes, I must inform you that while you have provided the title of the paper—Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations—and detailed formatting instructions, the actual text or content of this scientific paper was not included.

The tables provided relate to Per-block test accuracy of LR and MLP probes for Llama-3.2-3B-Instruct, Qwen-3..., which is a different topic from detecting refusal signals.

To fulfill your request—which requires extracting 450 to 600 words of highly structured, quoted information directly from the source material—I need the full text of Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations.

Please provide the body text of the paper, and I will immediately deliver the summary following your exact specifications:

  1. One short orienting paragraph (no header).

  2. 3 to 5 sections with bold headers (e.g., "How it works").

  3. Detailed paragraphs and bulleted/numbered lists using quoted key phrases from the paper, ensuring no commentary or external information is added.

Improvements for AI systems

The core scientific insight presented is that model knowledge and internal representations can be precisely measured and characterized at every intermediate layer (transformer block) using linear (LR) and Multi-Layer Perceptron (MLP) probes. This capability moves beyond simple end-to-end evaluation.

Here are the specific, high-impact improvements that can be implemented in AI systems:


Improvement: Instead of treating the LLM as a monolithic black box, integrate dedicated Knowledge Fingerprinting Modules (KFMs) at regular intervals across the transformer architecture. These modules are trained simultaneously with the primary task objective, using linear or MLP probes to explicitly map and quantify specific types of knowledge (e.g., factual recall, syntactic structure recognition, causal relationship identification) within the intermediate activations D(l).

What the Improved AI System Can Do:

  • Diagnostic Debugging: When the system fails (e.g., hallucination, logical inconsistency), the KFM can pinpoint which specific block (l) and what type of knowledge (e.g., causality) was incorrectly processed or lost, providing immediate diagnostic feedback to human operators or fine-tuning loops.

  • Attribution and Explainability (XAI): The system can generate a detailed Knowledge Attestation Report for its output, listing the specific transformer blocks that contributed positively to the answer and identifying any blocks where the representation was weak or misleading. This elevates explainability from post-hoc analysis to an intrinsic part of the inference process.

  • Targeted Retraining: Instead of expensive full-model fine-tuning, weaknesses can be isolated. If KFM detects a persistent dip in temporal reasoning accuracy at block 35, only the parameters related to that specific knowledge representation can be targeted for minimal-data Continual Learning updates.

Sources

Related papers