Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models

arXiv:2610.10625 · cs.CR, cs.LG · Submitted 2026-10-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Safe at One Loop, Risky at Another".

Elias: The gist The safety of a LoopLM cannot be inferred from a single recurrent depth, motivating safety alignment across recurrent computation.

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: We're moving into this next part where we look at what's actually happening in this paper, 'Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models'. This paper is fundamentally concerned with the safety of Looped Language Models which scale up using shared parameters.

Elias: The central thesis here is that since a single LoopLM has outputs across multiple inference depths, we must investigate whether safety is maintained throughout all those recurrent steps, because prior tests just didn't show us this robustness yet.

Priya: They point out the two key issues they found: first, safety can vary depending on the inference depth you check for a specific jailbreak input.

Nadia: And second, attacks built against one source depth can jump over and work on other depths, creating multiple attack sources for any target depth you pick.

Elias: It matters because this means current safety checks aren't enough; we need a way to align safety across these different recurrent depths instead of just hoping the final output is safe.

Priya: So they are setting out to study this variation and alignment across recurrent depths, which is something that hasn't been much explored in the field yet.

Conclusion: Nadia: Looking at the title, 'Safe at One Loop, Risky at Another', it really captures the tension they found: safety isn't guaranteed just because you check one level of computation; it depends entirely on which depth you are looking at.

Elias: And the authors, Yi Wang and others, they are essentially pointing out that scaling up with Looped Language Models introduces a new complexity where we can't treat every recurrent step as having the same safety guarantee.

Priya: What this means for me is that future work needs to focus on building systems that inherently understand this multi-depth nature of the computation, not just trying to patch it after training.

Nadia: This work motivates us to think about how we align safety across recurrent computation rather than just focusing on a single output depth for these powerful models.

Yi Wang, Xiuyuan Qi, Dongqi Han, Dongsheng Li, Wenjie Wang

ShanghaiTech University · Microsoft Research Asia

cs.CR, cs.LG

Submitted: 2026-10-07

Updated: 2026-10-07

Comments: 25 pages, 6 figures. Submitted to ICLR 2027

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

The gist: The gist The safety of a LoopLM cannot be inferred from a single recurrent depth, motivating safety alignment across recurrent computation.

Key concepts

Looped Language Models (LoopLMs)
These are language models where the output is exposed at multiple different inference depths due to their recurrent structure. This means a single query can yield different results depending on how deep the model processes it, which complicates safety evaluation.
Safety Variation Across Inference Depths
This challenge shows that an attack's success rate changes depending on which inference depth is targeted. Attacks might succeed against one depth but fail against another, making safety unpredictable across the model's recurrent steps.
CrossDepth Jailbreak Transfer
This occurs when a jailbreak crafted to bypass safety at one inference depth can also successfully bypass safety at other depths. This means attackers have multiple ways to exploit vulnerabilities within the model's various recurrent layers.

Terminology

Summary

The gist The safety of a LoopLM cannot be inferred from a single recurrent depth, motivating safety alignment across recurrent computation.

Safety Challenges in Looped Language Models

Looped Language Models (LoopLMs) expose outputs at multiple inference depths because a single LoopLM exposes outputs at multiple inference depths across different recurrent steps. Prior evaluations suggest that deeper recurrence can improve safety on harmful queries, but robustness under jailbreak attacks remains unclear. The evaluation reveals two safety challenges: Safety Variation Across Inference Depths and CrossDepth Jailbreak Transfer. First, Safety Variation Across Inference Depths shows that attack success rate (ASR) decreases with depth on original harmful queries but can increase under jailbreak attacks. Second, CrossDepth Jailbreak Transfer indicates that jailbreaks constructed against one depth can transfer to others, allowing multiple attack sources for each target depth.

SafeBridge Architecture

To address these gaps, SafeBridge is introduced as a preference optimization method for cross-depth safety alignment in LoopLMs. Architecturally, SafeBridge combines lightweight depth specific modulation with a selective state bridge from an earlier recurrent depth to later depths. Depth Specific Modulation allows each supervised depth to adjust the shared computation with its own parameters, where the input at recurrent depth d is defined by Eq. 1. The Selective State Bridging introduces a gated bridge from source depth sc to target depths T ⊆ T: κ(d)i = (tanh(vd)⊙ sg h(sc)i, d ∈ T, 0, d /∈ T, h(d)i = Norm z(d)i,K + κ(d)i.

Training Objective

The training objective for SafeBridge combines pairwise preferences with response level safety labels, emphasizing weaker depths with uniform supervision across depths. The final complete objective is LSafeBridge = 1/2(E + U) + γF, where E is the weak-exit aggregation and U is the uniform supervision. The components include depth specific modulation u(d)i,k = Fθk z(d)i,k-1, d ∈ D.

Experimental Results

SafeBridge substantially reduces ASR for both matched-depth and cross-depth attacks across model scales and benchmarks. For Ouro-1.4B on Ouro-2.6B with DPO, SafeBridge reduces the attack ASR from 44.93% to 9.85%. Furthermore, SafeBridge reduces the pooled source ASR for all three attack methods at both model scales, demonstrating robustness when attackers use candidates constructed against multiple source depths.

General Utility and Over-Refusal

SafeBridge improves general utility over vanilla models while maintaining comparable over-refusal behavior to SafeDPO. The ablation studies show that removing the response safety loss causes the largest increase in mean ASR, while removing modulation raises worst setting ASR from 24.00% to 59.00%.

Conclusion

SafeBridge addresses cross-depth safety through depth specific control of shared recurrent computation and joint alignment of depth indexed policies. Its components address complementary aspects of cross-depth safety alignment. These findings highlight the importance of evaluating and aligning safety across recurrent depths and motivate future safety alignment methods that explicitly account for recurrent computation.>

The paper is a comprehensive evaluation of Looped Language Models (LoopLMs) to determine if safety is preserved across their repeated recurrent steps. It systematically evaluates two key challenges: Safety Variation Across Inference Depths and CrossDepth Jailbreak Transfer. The proposed SafeBridge method successfully mitigates these issues by combining depth specific modulation and selective state bridging with response safety supervision. This method substantially reduces attack success for both matched-depth and cross-depth attacks while maintaining comparable over-refusal behavior and improving general utility. The findings motivate the necessity of safety alignment across recurrent computation rather than relying on a single output depth.

How it works

The methodology involves conducting a comprehensive safety evaluation of Looped Language Models under jailbreak attacks across multiple model scales, benchmarks, attack methods, and recurrent depths. The evaluation reveals two safety challenges: Safety Variation Across Inference Depths and CrossDepth Jailbreak Transfer. <ref:2610.

Improvements for AI systems

  1. Improved Safety Alignment Across Recurrent Depths: SafeBridge introduces depth specific modulation allowing each supervised depth to adjust shared computation with its own parameters, addressing how an update targeting one depth inevitably changes the computation at other depths.

  2. Enhanced State Information Flow: The selective state bridging mechanism allows later depths to access earlier states from a relatively safer source, defined by the equation: κ(d)i = (tanh(vd) ⊙ sg h(sc)i, d ∈ T, 0, d /∈ T, h(d)i = Norm z(d)i,K + κ(d)i.

  3. Robust Safety Supervision: The training objective incorporates response safety labels to promote safe responses and suppress unsafe ones via the term b j d = softplus1 − q j d, which is combined with preference losses in the final objective: LSafeBridge = 1/2 (E + U) + γF.

  4. Adaptive Loss Emphasis: The model uses a normalized smooth maximum, SMτ (w) = τ log 1/D X d∈D e w d / τ, to create an adaptive objective that emphasizes weaker exits: The smooth maximum gives larger gradient weights to larger losses, emphasizing weaker exits.

  5. Cross-Depth Consistency Enforcement: A cross-depth consistency objective, defined by the term ϕ j d→d+1 = τ softplus sg(q j d) − q j d+1 / τ, is included to discourage safety degradation as recurrent computation proceeds.

Sources

Related papers