SafeDepth: Safety-Aware Token-Level Adaptive Computation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "SafeDepth: Safety-Aware Token-Level Adaptive Computation".
Nadia: Token-level adaptive computation allows different tokens to execute different subsets of Transformer layers, but existing methods show that these execution choices negatively impact model safety.
Elias: First, who's behind it and why it matters.
Paper summary: Nadia: SafeDepth introduces this lightweight framework that learns to selectively execute or skip layers for each token based on their representations and the current inference phase, aiming to improve both safety and computational efficiency while keeping the pre-trained backbone frozen. That's a major claim because we're trying to use these models more broadly.
Elias: The core thesis is that existing token-level adaptive models show higher harmful response rates than their reference backbones, and SafeDepth aims to fix that by learning which computations support safe responses and bypassing those that contribute to unsafe generation. That addresses the fundamental issue of performance versus safety in these architectures.
Priya: It sounds like the framework is designed not just for speed, but specifically to preserve the safety guarantees of a larger model while reducing its operational footprint during inference tasks. I'm interested in how this selective execution actually maps back to preserving task capability on benign tasks, which is crucial for real-world deployment.
Nadia: They claim that SafeDepth balances generation quality and computational cost using a router that selects layer execution based on token representations, layer position, and the inference phase. This joint training aims to find paths that are both efficient and safe without needing an external safety judge during the actual inference run.
Elias: Furthermore, they use a two-stage training process; Phase I establishes these computation paths using a loss function involving cross-entropy for prediction accuracy and KL divergence to align with a frozen teacher's distribution. Then, Phase II refines these paths using safety feedback from complete responses through coarse-to-fine route localization.
Priya: The refinement stage sounds interesting because it acknowledges that unsafe answers don't directly tell the system which routing decisions need changing, so they use a sophisticated method to locate those safety-sensitive computations before switching execution paths. That suggests a deeper understanding of the internal state required for safety.
Nadia: Ultimately, SafeDepth claims that its approach leads to tangible improvements across various benchmarks; for instance, it reduces unsafe response rates on HarmBench-HJ from fifty-five point four four percent down to twenty-five point two five percent, which is a significant absolute decrease. That's the kind of concrete evidence we look for when evaluating new methods.
Elias: The paper also notes that compared to models like FlexiDepth, SafeDepth lowers the XSTest false-refusal rate from twelve point zero percent to four point zero percent, showing a clear improvement in how it handles refusal scenarios without sacrificing too much performance quality on tasks like GSM8K and HumanEval+.
Conclusion: Nadia: So, thinking about the title, SafeDepth really speaks to making the adaptation process itself safety-aware, moving beyond just optimizing for speed or accuracy on a token level. The authors managed to create a mechanism that learns to selectively execute computation paths based on context and phase without needing an external judge during inference.
Elias: The implications here are significant because it suggests we can prune the computational load of large models while maintaining safety properties that were previously only guaranteed by running every layer fully. This could make deploying very complex reasoning models to resource-constrained environments much more viable in practice.
Priya: What I see is that this research shifts the focus from just training a model to optimizing its dynamic inference behavior, which has direct implications for privacy because we are potentially reducing the computational footprint during processing sensitive data. It's about making safety an intrinsic part of the efficiency mechanism itself.
Nadia: And I think the authors' work on refining those paths using complete-response interventions really highlights a necessary step: you can't just prune blindly; you have to use feedback from full, safe responses to guide those pruning decisions effectively. That’s a practical insight for anyone trying to deploy these techniques.
Elias: It confirms that the relationship between computation and safety is complex and dependent on layer-specific interactions, not just a simple measure of executed layers. We need to keep testing those assumptions about how different components interact under adversarial conditions.
Priya: And from a measurement perspective, the validation showing task capability is maintained or even improved on GSM8K and HumanEval+ gives us confidence that this efficiency gain doesn't come at the cost of core reasoning ability, which is what we need to measure rigorously.
Nadia: It sounds like SafeDepth provides a much more nuanced toolkit for engineers looking to deploy these kinds of adaptive architectures responsibly. It moves the conversation toward designing models where safety and efficiency are jointly optimized objectives from the very beginning.
Nizhang Li, Ian G. Harris
Macau University of Science and Technology · University of California, Irvine
cs.CR
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/evalplus/evalplus
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 86/100
The gist: Token-level adaptive computation allows different tokens to execute different subsets of Transformer layers, but existing methods show that these execution choices negatively impact model safety.
Key concepts
- Token-level Adaptive Computation
- This technique allows different parts of the Transformer model (tokens) to run different sets of layers during inference. SafeDepth learns which layers each token should skip or execute based on what it sees and when it is generating an answer, aiming for better efficiency without sacrificing quality.
- SafeDepth Framework
- This is a lightweight system that uses a router and adapters to decide whether to run or skip layers for each token. It is trained in two stages: first to balance generation quality with cost, and second using safety feedback to refine the paths, ensuring the final model remains safe.
- Coarse-to-fine Route Localization
- This is a refinement technique used during training where the system changes routing decisions for all tokens and regenerates answers. It helps narrow down the search space for optimal layer skipping by prioritizing combinations of changes that lead to successful safety improvements.
Terminology
Summary
Token-level adaptive computation allows different tokens to execute different subsets of Transformer layers, but existing methods show that these execution choices negatively impact model safety. The research introduces SafeDepth, a lightweight framework that learns to selectively execute or skip layers for each token based on token representations and inference phase, aiming to improve both safety and computational efficiency while preserving performance on benign tasks.
The gist
SafeDepth is a safety-aware token-level adaptive reasoning framework that learns execution and skipping decisions while keeping the pretrained backbone frozen.
Motivation and Findings
The research investigates whether selective layer execution affects model safety, finding that existing token-level adaptive models exhibit higher harmful response rates than their reference backbones. The study reveals several key dependencies:
-
Safety effects depend on which layers are skipped and whether skipping occurs during prompt processing or answer generation.
-
Recognizing harmful requests and producing refusals depend on different layer-wise computations, indicating that safety is not solely dependent on the amount of computation performed.
-
Restoring one skipped layer at a time can have varying effects;
Restoring one skipped layer at a time reduces unsafe-response rates for some layers but can increase them for others.
SafeDepth Architecture and Training
SafeDepth utilizes two-stage training to establish computation paths that balance generation quality and computational cost. The framework consists of a lightweight router and adapters, which are jointly trained to reduce computation while preserving generation.
-
The router selects execution based on
token representations, layer position, and inference phase.
-
When a layer is skipped, an adapter maps the incoming hidden state to a
same-dimensional replacement for the next layer.
-
Phase I jointly trains the router and adapters using a loss function that includes cross-entropy for prediction accuracy and KL divergence to ensure agreement with a frozen, fully executed teacher’s distribution.
Safety-Guided Refinement
The second stage of training refines these paths using safety feedback from complete responses, addressing the challenge that unsafe answers do not directly reveal necessary routing changes. This is achieved through Coarse-to-fine route localization.
-
The process involves changing the routing actions of one layer for all prompt tokens and regenerating the complete answer to define a smaller search region.
-
The system evaluates
bounded combinations
of changes, retaining successful groups rather than treating members as independently effective. -
The optimization objective LII prioritizes successful changes:
The ranking term gives successful changes higher priority than the other tested candidates.
Experimental Validation and Trade-offs
Experiments on Llama-3-8B-Instruct validate SafeDepth's effectiveness against baselines like FlexiDepth and DiffSkip. The results demonstrate a favorable trade-off between task performance, safety, and computational cost.
-
SafeDepth reduces unsafe-response rates across all five harmful-request benchmarks, with the largest absolute decrease on HarmBench-HJ (from 55.44% to 25.25%).
-
Compared with FlexiDepth, SafeDepth lowers the XSTest false-refusal rate from 12.0% to 4.0%.
-
On WildJailbreak, SafeDepth reduces the unsafe-response rate from FlexiDepth’s 41.67% to 30.21% while lowering complete-response matrix work by 24.5%.
-
Task capability is maintained or improved: SafeDepth improves on GSM8K and HumanEval+, achieving a
better safety–efficiency trade-off
than FlexiDepth, which sacrifices safety for lower cost in some areas.
Key Analytical Insights
The analysis establishes four critical findings regarding the impact of adaptive computation on safety:
-
Adaptive computation does not preserve backbone safety;
All token-level adaptive reasoning models in our study exhibit higher unsafe-response rates than their corresponding backbones.
-
Prefill and decoding have different safety effects;
Restoring prefill computation reduces unsafe-response rates more consistently than restoring it during decoding.
-
Safety is layer-dependent:
Safety is therefore not a function of executed-layer count alone,
as restoring computation at different layers produces distinct safety outcomes. -
Refusal depends on the execution path:
Recognizing harmful intent does not ensure a safe response,
as later execution decisions can reverse the final safety outcome.
Ablation Study Summary
Ablation studies confirm that components like K/V preservation and skip adapters are crucial for maintaining task performance, as removing them leads to significant drops in exact match accuracy (e.g., GSM8K dropping from 66.64% to 0.83%). Furthermore, the Phase II refinement stage is necessary for safety improvement; removing it results in a higher unsafe-response rate (46.80%) compared to the refined model (30.21%).
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed the provided paper, SAFEDEPTH: SAFETY-AWARE TOKEN-LEVEL ADAPTIVE COMPUTATION.
The core contribution is a lightweight framework that learns token-level execution policies (execute or skip) within Transformer layers to simultaneously improve safety and efficiency.
Here are the specific improvements and capabilities this system can provide to AI systems:
The improved AI system, powered by SafeDepth, will possess the following capabilities:
-
Inference of Safety-Aware Computation Paths: The system will dynamically decide, on a token-by-token basis and layer-by-layer, whether to execute or skip a Transformer block based on the specific input token's representation and its position within the model (prefill vs. generation phase).
-
Safety Refinement via Response Feedback: Unlike previous methods that only optimize for generation quality, SafeDepth utilizes a two-stage training process where Phase II refines the routing decisions using safety feedback from complete responses. This allows the model to learn which specific layer-wise computations are responsible for producing unsafe outputs versus safe ones.
-
Targeted Computation for Harmful Intent Recognition: The system can be tuned to retain computation in specific layers known to be critical for recognizing harmful intent (as identified by the safety-sensitive computation analysis in Section 3), ensuring that danger signals are processed effectively before the final response is generated.
-
Phase-Aware Routing: The system distinguishes between prefill (prompt processing) and decoding (answer generation). It can be explicitly configured to restore specific layers during prefill to enhance prompt understanding or, conversely, skip layers during decoding if they are correlated with harmful output generation.
-
Improved Safety vs. Efficiency Trade-off: The system achieves a superior trade-off compared to existing token-level adaptive models (like FlexiDepth). It can produce responses that are significantly less likely to be harmful across multiple benchmarks (e.g., 25% unsafe response rate on HarmBench-HJ) while reducing total computational cost by up to 24.5% on WildJailbreak compared to full-depth inference.
-
Enhanced Benign Response Quality: Crucially, the improved system does not simply increase refusal rates for benign requests. By learning safer routing policies, it can actually improve the quality of answers for benign tasks (e.g., GSM8K accuracy increased from 65.66% to 66.64%), demonstrating that safety-aware computation is compatible with stronger task performance on useful generation tasks.
-
Causal Understanding of Safety Failure: The framework provides insights into the causal link between recognizing harmful intent in intermediate representations and the final behavioral output, allowing researchers to identify precisely which subsequent computational steps are responsible for
turning
a risk signal into an unsafe response.
Sources
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference
- AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference
- Routing-Aware Safety Alignment for Mixture-of-Experts Models
- Let's Verify Step by Step
- Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
- Learning to Skip for Language Modeling
- LLMs Encode Harmfulness and Refusal Separately
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs