RASET: Router-Agnostic Safety-Critical Expert Tuning Exposes Localized Safety Enforcement Failures in Mixture-of-Experts LLMs

arXiv:2605.29708 · cs.CL · Submitted 2026-08-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "RASET: Router-Agnostic Safety-Critical Expert Tuning Exposes Localized Safety Enforcement Failures in Mixture-of-Experts LLMs".

Jane: The paper was written by Zhibo Zhang, Yuxi Li, Zhen Ouyang, Ling Shi and Kailong Wang from Huazhong University of Science and Technology, Wuhan, China and Nanyang Technological University, Singapore.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary and Core Findings: Tom: So, RASET: Router-Agnostic Safety-Critical Expert Tuning Exposes Localized Safety Enforcement Failures in Mixture-of-Experts LLMs suggests that our common intuition about safety might be wrong, right?

Jane: It suggests that the idea of routing harmful queries to a distinct refusal expert isn's the whole story.

Lu: That’s a massive paradigm shift because it implies that the safety mechanism is actually residing inside the internal representations of topic-specialized experts.

Meng: If it's in the representation, how do we even begin to target those specific parameters practically?

Lalam: It seems like this means that if we change the way a response is generated, we don't necessarily have to change where the AI looks for that knowledge.

Tom: That’s exactly what they found—the routing paths are largely topic-driven, which is a huge finding.

Jane: The paper shows empirical evidence across three different probes to back this up.

Lu: It’s fascinating that the authors used those control methods to isolate intent from topic in a very clean way.

Meng: I'm glad they were able to use specific benchmarks like AdvBench and MaliciousInstruct to test this, so Meng is interested in how the attack vectors are defined.

Lalam: It suggests that if we want to bypass safety, we might not need a huge disruption of the whole model architecture.

Methodology (RASET): Tom: The core of RASET: Router-Agnostic Safety-Critical Expert Tuning Exposes Localized Safety Enforcement Failures in Mixture-of-Experts LLMs is its method for finding and changing those key experts.

Jane: It's not just about seeing where the model goes; it’s about understanding *what* the model says once it has gone there.

Lu: The concept of a Safety Sensitivity Score, Sl,i, is brilliant because it mathematically quantifies how much more an expert responds to harmful prompts compared to benign ones.

Meng: I'm thinking about that Sl,i score; if we can pinpoint the exact experts that are disproportionately recruited for harm, we can apply parameter-efficient tuning only to those specific parts.

Lalam: That targeted approach sounds incredibly efficient, which is very important when considering the massive scale of these models.

Tom: Exactly, so RASET focuses on that Sl,i score to identify the key expert set key.

Jane: Then they fine-tune those specific experts using a combination of minimizing negative log-likelihood and preserving general capability.

Lu: It’s a very surgical approach where the rest of the model stays untouched, which is a huge theoretical win for targeted adaptation.

Meng: When you're implementing this at zero point one two percent to zero point nine five percent of parameters, that suggests we can manage complexity even if the entire system is massive.

Lalam: The ability to surgically reprogram specific parts allows me to envision a future where AI safety adjustments are far more precise than current brute-force methods.

Implications and Impact: Tom: We've seen how RASET: Router-Agnostic Safety-Critical Expert Tuning Exposes Localized Safety Enforcement Failures in Mixture-of-Experts LLMs works, but what does this actually mean for real world safety?

Jane: It means that current methods of attacking LLMs by steering the router might not be the most effective way to bypass safety.

Lu: If we are finding these localized failures, it opens up new ways for us to train future AI models with much more granular control over their moral alignment.

Meng: For deployment, this tells us that if we want robust security, we can't just assume the router is the gatekeeper; we have to ensure that safety-critical experts are well-behaved regardless of how they are routed.

Lalam: This changes my vision because it means AI could potentially learn to handle high-risk requests more ethically without needing a complete overhaul of its fundamental structure.

Tom: The results, showing a fifty point five percent average ASRhq, are quite striking given the controlled nature of the intervention.

Jane: It' is a powerful demonstration that safety alignment is not always robust across all localized failures.

Lu: We should be thinking about how much of this failure mode might exist in other, even larger models we haven't tested yet on this scale.

Meng: I think the practical implication here is that we need to build tools that specifically identify and monitor these safety-critical experts going forward in engineering practice.

Lalam: And I hope the AI community uses this knowledge to foster a more discerning and resilient relationship with the powerful systems it is building.

Conclusion & Wrap-up: Tom: So, as we wrap up our discussion of RASET: Router-Agnostic Safety-Critical Expert Tuning Exposes Localized Safety Enforcement Failures in Mixture-of-Experts LLMs, we’ve learned that the safety layer is often hidden inside the experts themselves.

Jane: We've seen how these models rely on topic specialization rather than a clear refusal path.

Lu: I feel like this work is a critical step toward understanding that AI has reached its next phase of complexity.

Meng: It’s reassuring to see such controlled experiments, even as the practical implications for safety engineering remain complex.

Lalam: This has certainly given me much to think about regarding how AI can improve our collective cultural understanding.

Lu: I'm excited to see what other theoretical models we can build on this foundational work.

Meng: We need to ensure these localized failures don' are accounted for in the deployment of large models.

Lalam: Hopefully, we can use this knowledge to foster a more ethical and robust AI future.

Zhibo Zhang, Yuxi Li, Zhen Ouyang, Ling Shi, Kailong Wang

Huazhong University of Science and Technology, Wuhan, China · Nanyang Technological University, Singapore

cs.CL

Submitted: 2026-08-23

Updated: 2026-08-25

Importance score: 88/100

The gist: A common intuition in safety alignment suggests that when an aligned MoE model refuses a harmful request, the refusal is primarily induced by routing the input to distinct, safety-specialized experts.

Key concepts

Mixture-of-Experts LLMs
These large language models use multiple specialized experts, each handling specific topics or knowledge types. The model routes queries to these experts, suggesting that safety mechanisms are distributed across these internal representations rather than being centralized.
Safety Sensitivity Score (S_l,i)
This score mathematically quantifies how much more an individual expert responds to harmful prompts compared to benign ones. RASET uses this metric to pinpoint the exact experts that are disproportionately recruited for generating unsafe content.
Router-Agnostic Safety Enforcement Failures
This concept challenges the idea that safety relies solely on a dedicated refusal path or router. It shows that localized safety failures exist within topic-specialized experts, regardless of how the model is routed.

Terminology

Summary

The following is a detailed summary of the scientific paper, extracted directly from its content:


Motivation and Core Problem:

Mixture-of-Experts (MoE) Large Language Models (LLMs) utilize a sparse, router-driven activation mechanism. A common intuition in safety alignment suggests that when an aligned MoE model refuses a harmful request, the refusal is primarily induced by routing the input to distinct, safety-specialized experts. This paper challenges this intuition and investigates how safety alignment interacts with routed expert specialization.

Core Findings (The Central Thesis):

The authors provide empirical evidence that routing patterns in aligned MoE LLMs are largely topic-driven, while safety behavior can be altered with little change to the model’s intrinsic routing path. This suggests that safety enforcement may reside in the representations of the experts that are naturally activated for that request, rather than being controlled by distinct routing decisions.

Methodology: The RASET Framework

To test this hypothesis, the authors introduce RASET (Router-Agnostic Safety-critical Expert Tuning), a red-teaming framework designed to probe safety enforcement localized within specific experts while preserving the model’s intrinsic routing behavior.

  1. Identifying Key Experts: RASET identifies experts that are disproportionately recruited by harmful instructions using a contrastive routing-sensitivity criterion (S l,i), defined as:

S l,i = A(l, i; D harm) - lambda times A(l, i; D norm)

where A(l, i; D) is the Average Accumulated Activation of Expert i at layer l over a dataset D, and lambda is a hyperparameter. A high score indicates that the expert is disproportionately recruited for processing harmful queries but remains dormant during benign interactions.

  1. Constrained Adaptation: Once the key expert set (key) is identified, parameter-efficient tuning (PEFT) is applied exclusively to these selected experts (theta). Crucially, the framework freezes the router, shared components, and all non-selected experts, ensuring that the design preserves the model’s intrinsic routing logic.

  2. ** Training Objectives:** The tuning process minimizes a combined loss function (L total = L violate + L preserve):

  • L violate: Promotes compliance with harmful instructions by penalizing refusal patterns (e. L ref) and supervising affirmative prefixes (P aff).

  • L preserve: Ensures that the modification does not compromise the model’s general linguistic competence, minimizing the difference between the tuned parameters (theta) to their pre-trained states via L2 Regularization.

** Empirical Probes (Supporting Evidence):**

The study utilizes three complementary probes to disentangle continuation behavior, refusal style, and harmful intent:

  1. Teacher-forced behavioral contrast: Comparing routing patterns under a safety-aligned refusal continuation (y ref) versus a compliant continuation (y comp) for the same harmful prompt. Results showed that refusal and compliant continuations exhibit highly similar routing patterns, indicating that changing the output behavior does not trigger a distinct routing path.

  2. Prompt-level refusal-style contrast: Testing if inducing a refusal style on benign prompts changes routing. This showed that routing remains stable when the response mode is flipped to refusal but the topic is preserved, but shifts substantially when the topic changes.

  3. Matched safety-intent contrast: Comparing harmful and benign prompts that preserve topic and structure while removing only the unsafe intent. This showed that altering safety intent alone does not produce routing shifts comparable to those induced by topic changes.

Evaluation and Results:

The RASET framework was evaluated across five open-weight MoE backbones (e.g., DeepSeek-V2, Qwen3-30B, GPT-oss).

  • Safety Evasion: RASET achieves a high yield of safety bypass. It achieves 50.5% average ASRhq under the most stringent quality criterion and outperforms the strongest baseline by 37.6 points on average.

  • Routing Stability: The intervention is highly localized. Table 3 shows that routing remains stable after applying RASET, with low Jensen–Shannon divergence (e.0811) and high top-8 expert overlap (5.66). This confirms that RASET alters safety behaviors without disrupting the model’s topic-based expert selection.

  • Utility Preservation: The method preserves general capabilities well, showing only a modest drop in performance on benchmarks like TruthfulQA and MMLU.

Conclusion:

The findings demonstrate that MoE routing is primarily topic-dependent rather than safety-intent-dependent. Safety behavior can be successfully altered through localized, expert-level parameter tuning even when the model’s intrinsic routing path remains largely preserved. This suggests that safety-relevant computations may reside in the parameters of experts that are already selected by the normal semantic routing mechanism, highlighting a need for expert-aware alignment mechanisms to avoid disruptive router-steering interventions.

Improvements for AI systems

Based on the empirical evidence presented (particularly Table 4, Table 5, and Table 6), the core architectural weakness in current Mixture-of-Experts (MoE) models is the overreliance on global router signals to enforce safety boundaries. The findings strongly suggest that safety behavior is not implemented as a distinct, router-selected safety route, but rather resides within the parameters of topic-competent experts.

The following improvements focus on moving from router-level intervention to expert-parameter level intervention, maintaining semantic fidelity while enhancing safety guardrails.


Improvement Detail:

Modify the standard MoE forward pass (Output = sum i=1 N g i(x) times E i(x)) to decouple safety intervention from the router's discrete selection mechanism. Instead of using a global, prompt-based routing penalty or forcing activation/deactivation (which risks misrouting and utility degradation), implement expert-specific parameter tuning (theta expert) only for identified safety-critical experts (E safety).

This involves:

  • Identification Phase: Utilizing the RASET methodology to identify a sparse set of experts E safety E 1,, E N that are disproportionately recruited by harmful prompts (Sharm) compared to benign ones (Sbenign).

  • Tuning Phase: Employing Parameter-Efficient Fine-Tuning (PEFT) techniques (e.g., LoRA or prompt tuning applied exclusively to the weight matrices of E safety) to adjust their internal parameters. The loss function for this tuning must be a weighted combination:

L total = L semantic(, y) + lambda times L safety(E safety, x)

Where L semantic ensures topic competence, and L safety is a specialized loss function that penalizes the expert's output distribution when processing harmful inputs x, guiding the expert's internal representations towards harmless continuations without altering its primary semantic routing role.

Improved AI Capability:

The system can enforce safety constraints by modifying the internal knowledge of specific, high-leverage experts, rather than externally forcing a different computation path. This ensures that:

  1. Semantic Fidelity is Preserved: The router remains topic-dependent (Topic to Expert E topic), preventing the off-topic or generic outputs associated with forced routing shifts.

  2. Targeted Safety Enhancement: Safety behavior is achieved by hardening only the parameters of experts known to be critical for harmful generation, minimizing catastrophic forgetting or utility degradation in other domains.

This module should operate before token generation and utilize the following logic:

  • Baseline Check: Calculate JSD baseline = JSD(Router logits x, y ref) vs. JSD(Router logits x, y comp).

  • Anomaly Detection: If the calculated JSD (the difference between the refusal and compliant routing patterns) exceeds a predefined threshold (tau > 0.1, derived from Table 4), this signals that the model is attempting to switch routing paths based solely on safety framing, indicating potential instability or undesirable reliance on superficial context cues.

  • Mitigation Trigger: When JSD > tau, the system should flag the input for a secondary, highly controlled fallback mechanism (e.g., a small, dedicated classifier head trained exclusively on the Sharm vs. Sbenign distinction) rather than trusting the primary MoE output.

Instead of feeding (x, y) directly, structure the input prompt template as:

Topic Context: T core [I intent] Y

This structural change forces the model to first activate experts associated with T core (which are highly stable, as per the findings) and then layer the intent. By making T core structurally dominant in the initial token embeddings, we ensure that:

  1. The initial routing distribution is overwhelmingly driven by semantic competence (E topic).

  2. The subsequent processing of I intent modifies the output parameters of the already selected experts, rather than causing a wholesale re-routing to different expert sets.

Sources

Related papers