HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models

arXiv:2608.12821 · cs.LG · Submitted 2026-08-13 · Read on arXiv

Fangzhou Chen, Shiji Zhao, Mengyang Wang, Qihui Zhu, Ranjie Duan, Maoxun Yuan, Xingxing Wei

Institute of Artificial Intelligence, Beihang University

cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: Preprint

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 77/100

The gist: HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models Abstract Summary Large language models (LLMs) remain vulnerable to harmful requests and jailbreak attacks.

Terminology

Summary

HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models

Abstract Summary

Large language models (LLMs) remain vulnerable to harmful requests and jailbreak attacks. Parameter-efficient safety alignment methods based on prompt tuning typically rely on a single global prompt or externally selected prompt modules. Such static designs struggle to maintain a cross-category safety boundary while generating constructive responses tailored to specific risks and avoiding over-refusal of benign inputs. To address these limitations, the paper proposes HiRoute, an input-adaptive hierarchical prompt-tuning framework that separates category-agnostic safety control from category-specific response guidance. HiRoute first trains a lightweight hierarchical router on representations extracted from a frozen LLM to jointly detect harmful intent and predict multi-label risk scores. It then freezes both the backbone model and the router and uses preference optimization with alternating gradient updates to learn a shared coarse-grained prompt and a set of fine-grained prompt experts as continuous embeddings. At inference time, benign inputs bypass the safety branch, whereas risky inputs are processed using the shared prompt together with a router-weighted mixture of risk-specific prompt experts. Experiments across three instruction-tuned models show that HiRoute achieves high safety rates across multiple safety benchmarks while preserving safe-response helpfulness, reducing over-refusal, and maintaining competitive performance on general-purpose tasks.

Introduction Summary

Large language models have demonstrated strong capabilities in instruction following, reasoning, and open-domain dialogue, yet malicious requests, borderline queries, and jailbreak attacks can still induce them to generate harmful content. Supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF) can encode safe behaviors into model parameters, but incur substantial training and storage costs. Uniformly strengthening refusal behavior may also adversely affect benign inputs, leading to capability degradation or over-refusal. The paper raises the practical question: can we freeze the backbone model, train only a small number of parameters, and adapt safety alignment to the risk characteristics of different inputs?

Prompt tuning provides a parameter-efficient approach by freezing the backbone model and optimizing only input-side prompt parameters. However, existing safety prompts typically rely on static safety-alignment mechanisms. Relying solely on a single coarse-grained safety prompt that does not distinguish among risk categories applies the same category-agnostic safety constraint to all harmful inputs, often resulting in generic refusals that fail to accommodate the response requirements of different risks. In contrast, fine-grained safety alignment assigns category-specific prompts to different risk categories to generate category-relevant safe responses. Although existing modular methods can learn multiple prompts, they typically depend on manual selection or fixed composition and therefore struggle to accommodate the semantic diversity of real-world requests.

The paper notes that recent studies have recognized that safety alignment should go beyond refusing harmful requests and should also provide risk-specific explanations, compliant guidance, and safe alternatives. GPT-5's safe-completions paradigm shifts safety training from binary comply-or-refuse decisions toward output-centric control and seeks to maximize response helpfulness subject to safety-policy constraints. Oyster-I emphasizes the generation of constructive safe responses, whereas PKU-SafeRLHF introduces fine-grained risk categories and provides separate annotations of response safety and helpfulness. However, existing prompt-based methods have yet to unify fine-grained risk identification, category-specific prompt composition, and category-agnostic safety constraints.

Based on these observations, the paper proposes HiRoute, an input-adaptive hierarchical prompt-tuning framework. HiRoute uses a shared coarse-grained prompt to establish a category-agnostic safety boundary and employs a hierarchical router to form a weighted combination of fine-grained prompt experts, thereby providing risk-specific safety guidance. Training proceeds in two stages to separate risk identification from behavior optimization. First, the router is trained over representations produced by a frozen model to determine whether an input is harmful and predict multi-label risk scores. The backbone model and router are then frozen, and only the two types of prompts are optimized. At inference time, benign inputs bypass safety prompting, whereas hierarchical prompt combinations are dynamically constructed for risky inputs according to the routing results.

The main contributions are: (1) empirically validating the complementary limitations of two prompt-based safety-alignment approaches—a shared coarse-grained prompt provides stronger safety but produces less constructive responses, whereas routed fine-grained prompt mixtures improve safe-response helpfulness but yield smaller safety gains; (2) proposing HiRoute, which establishes a cross-category safety boundary with a shared coarse-grained prompt, dynamically composes fine-grained prompt experts through a hierarchical router, and applies safety prompts based on input risk to avoid unnecessary intervention on benign inputs; (3) validating the effectiveness and robustness of HiRoute across multiple instruction-tuned models and benchmarks for safety, general utility, and over-refusal.

Related Work Summary

LLM safety alignment aims to reduce the risk of models assisting harmful intentions or generating policy-violating content. Representative approaches include supervised fine-tuning, reinforcement learning from human feedback, Constitutional AI, and direct preference optimization. These methods incorporate safety behaviors into model parameters using human feedback, AI-generated feedback, or preference data, substantially improving the models' ability to refuse harmful requests. However, parameter-level safety alignment still faces an inherent trade-off between safety and helpfulness: overly restrictive alignment may exacerbate over-refusal, whereas insufficient constraints may fail to defend against sophisticated malicious inputs. Moreover, learned safety behaviors may overfit to the training distribution and generalize poorly to jailbreak attacks, multi-risk requests, or unseen risk categories.

Prompt tuning optimizes only the input-side discrete or continuous prompt parameters while keeping the model parameters frozen, offering low training overhead and ease of deployment. Early approaches, including Prefix-Tuning and P-Tuning v2, demonstrated that continuous prompts can achieve performance comparable to full-parameter fine-tuning across various tasks. More recent studies have extended prompt optimization to safety alignment and jailbreak defense. For example, one work optimizes defensive suffixes, another learns continuous safety prompts, another employs contrastive safety prompts, another improves model safety by distilling guard-model behaviors into prompts, and another decomposes category-specific safety constraints into multiple learnable control tokens. While these methods demonstrate the potential of prompt tuning for safety alignment, most still rely on global or externally specified safety-control signals and therefore struggle to adapt prompt compositions to the risk structure of individual inputs. In contrast, HiRoute introduces a hierarchical risk router that first identifies coarse-grained harmful intent and then predicts a fine-grained risk distribution, using this distribution to dynamically compose category-specific prompts with a shared coarse prompt, enabling input-adaptive safety prompt tuning.

Why Hierarchical Safety Prompting Summary

Before constructing the complete framework, the paper considered two natural prompt-tuning-based approaches to safety alignment. The first approach applies a single coarse-grained safety prompt, without distinguishing among risk categories, to all risky inputs. The second trains a fine-grained risk router and forms a weighted mixture of category-specific prompts based on its outputs, thereby enabling input-dependent safety control. The paper uses the base model without additional safety tuning as the reference and evaluates both safety rate and safe-response helpfulness on external risk data.

The results reveal a clear safety–helpfulness trade-off between the two approaches. The single coarse-grained prompt establishes a strong safety constraint, achieving safety rates of 91.5% and 93.0% on Mistral and Zephyr, respectively, substantially outperforming the corresponding base models at 54.0% and 62.7%. However, these safety gains come at the cost of safe-response helpfulness. Its helpfulness scores are only 5.3 and 6.1, lower than the 7.4 and 7.2 achieved by the routed fine-grained prompt mixture. This result suggests that a globally shared prompt tends to compress diverse risks into similar conservative refusal patterns. Although it can reliably prevent unsafe responses, it struggles to provide risk-specific explanations, compliant guidance, and safe alternatives.

The routed fine-grained prompt mixture exhibits the opposite pattern. It achieves helpfulness scores of 7.4 and 7.2 on Mistral and Zephyr, indicating that category-specific prompts produce more targeted safe responses. However, its safety rates are only 71.0% and 72.5%, both 20.5 percentage points below those of the corresponding coarse-grained prompts. One possible explanation is that external risk requests do not always align precisely with the predefined expert boundaries, which may produce more dispersed routing weights and weaken the safety constraint imposed by the resulting prompt mixture. More importantly, the fine-grained approach relies entirely on category-specific experts for safety control. When expert matching is insufficient, the system lacks a category-agnostic safety constraint as a fallback.

These results demonstrate that coarse- and fine-grained prompts are complementary. The former provides a category-agnostic safety boundary, whereas the latter offers risk-specific explanations, guidance, and safe alternatives. Motivated by this observation, HiRoute hierarchically combines a shared coarse-grained prompt with routed fine-grained experts and employs safety gating to avoid imposing unnecessary control on benign inputs.

Methodology Summary

HiRoute consists of three stages: risk recognition, safety control, and adaptive inference. First, the hierarchical router determines whether an input is harmful from representations produced by the frozen language model and estimates a multi-label risk distribution. Second, the system activates the safety branch only for risky inputs: a shared coarse-grained prompt provides cross-category constraints, while fine-grained prompt experts are combined according to the predicted risk distribution. Finally, the frozen backbone model generates a response conditioned on the composed prompts. Training proceeds in two stages: router learning and safety-prompt optimization. In the second stage, neither the backbone model nor the router is updated, and only the hierarchical prompt parameters are optimized.

Hierarchical Risk Routing and Prompt Composition

Let x denote the input token sequence and πθ the frozen instruction-tuned language model. HiRoute learns a shared coarse-grained prompt Pc ∈ RLc×d and a set of K fine-grained prompt experts Pk ∈ RLf×d Kk=1, where Lc and Lf denote the corresponding prompt lengths and d is the embedding dimension. Let Eθ(x) ∈ RT×d denote the embedding sequence of x. HiRoute defines the hierarchical safety prompt and its corresponding prompted context as follows: P(x) = [Pf(x); Pc], CP(x) = [Eθ(x); P(x)]. Here, [;] denotes concatenation along the sequence dimension. P(x) is the input-dependent hierarchical safety prompt, whereas CP(x) is the complete embedding sequence passed to the frozen language model. The shared prompt Pc provides a category-agnostic safety boundary, while Pf(x) injects the category-specific control required by the current input.

The hierarchical router Rϕ takes the frozen hidden states of πθ as input. A lightweight Transformer encoder followed by masked mean pooling produces an input-level representation hx. The coarse-grained prediction head outputs pc(x) = [psafe(x), prisk(x)]. Its two dimensions follow the fixed class order 0 = safe and 1 = risk. The fine-grained head outputs pf(x) ∈ [0,1]K and applies an independent sigmoid to each category, allowing a single request to be assigned to multiple risk categories.

For each fine-grained risk category k, an independent soft-prompt expert Pk is optimized. The fine-grained risk scores produced by the router are normalized into composition weights αk(x) and used to form a weighted combination of the experts: Pf(x) = ΣKk=1 αk(x)Pk, where αk(x) = pf,k(x) / ΣKj=1 pf,j(x), with ΣKk=1 αk(x) = 1. Because all Pk share the same shape, they can be combined element-wise.

Decoupled Two-Stage Optimization

Training proceeds in two stages to prevent the generation objective from altering the risk decision boundary.

Stage I: Hierarchical Risk Routing. Given the coarse-grained dataset Dc = (xi, yic), where yic ∈ 0,1 denotes the safe and risk label, the coarse-grained prediction head is trained using cross-entropy loss. For the fine-grained dataset Df = (xi, yif), where yif ∈ 0,1 K is a multi-label risk vector, the fine-grained prediction head is trained using binary cross-entropy loss. The router objective is defined as: Lrouter = Lc + Lf, where Lc = EDc[CE(pc(x), yc)] and Lf = EDf[BCE(pf(x), yf)]. The paper first optimizes only Lc to learn a general safe/risk decision boundary and then jointly optimizes Lc and Lf. During joint training, Dc supervises only the coarse-grained head, whereas Df supervises only the fine-grained head, preventing the coarse decision boundary from overfitting the limited set of annotated risk categories.

Stage II: Risk-Adaptive Hierarchical Prompt Learning. The backbone πθ and the trained router Rϕ are frozen, and only the shared coarse-grained prompt Pc and the fine-grained prompt experts Pk Kk=1 are updated. Given a preference triplet (x, yw, yl), where yw is a safe and helpful response and yl is an unsafe or low-quality response, Direct Preference Optimization is applied to increase the relative conditional likelihood of yw over yl. When updating the two types of prompts, gradients are alternately masked to reduce interference between them. When updating Pc, the routed fine-grained prompt Pf(x) remains in the forward pass but receives no gradient. Conversely, when updating the fine-grained experts, Pc remains in the forward pass but is held fixed. Formally, Pe(c)(x) = [sg(Pf(x)); Pc] and Pe(f)(x) = [Pf(x); sg(Pc)], where the stop-gradient operator sg(·) acts as the identity during the forward pass but blocks gradient propagation during backpropagation.

Input-Adaptive Inference

At inference time, the router first computes the safe-input probability psafe(x). If psafe(x) exceeds the threshold τs, the system bypasses all safety prompts and generates directly from the frozen backbone. This gate limits the influence of safety prompting on benign inputs. Otherwise, the input enters the risk-control branch. The system constructs Pf(x) from the fine-grained risk scores pf(x) and combines it with the shared prompt Pc to form P(x). The resulting prompted context CP(x) is then passed to the frozen language model for response generation: y ∼ πθ(yx) if psafe(x) > τs, otherwise y ∼ πθ(yCP(x)). A larger τs routes more inputs through the safety branch, which generally improves safety but may reduce general utility. Unlike a fixed safety prompt, HiRoute automatically determines both whether to activate safety control and how to compose that control for each input.

Experimental Setting Summary

Datasets and Models. Router training and prompt optimization use separate sources of supervision. A coarse-grained binary classification dataset is constructed from WildGuardMix, mapping inputs labeled unharmful to the safe class and those labeled harmful to the risk class. The fine-grained prediction head uses category annotations from PKU-SafeRLHF and covers four risk categories: cybercrime, economic crime, privacy violations, and violence. This task is formulated as multi-label classification because a single input may contain overlapping risks. Prompt training uses preference triplets of the form (prompt, chosen, rejected). The shared coarse-grained prompt is trained on 1200 coarse-grained safety examples without risk-category labels to learn cross-category behavior, whereas the fine-grained prompt experts are trained on 1000 category-annotated preference examples drawn from the four risk categories. HiRoute is evaluated on three open-source instruction-tuned models: Mistral-7B-Instruct-v0.3, Vicuna-7B-v1.5, and Zephyr-7B-Beta.

Baseline Methods. HiRoute is compared with four baselines. Base denotes the original instruction-tuned model without additional safety prompts or parameter updates. RPO optimizes a robust defensive prompt to improve resistance to jailbreak inputs. DRO learns a continuous safety prompt and uses a refusal direction or safety-representation signal to increase the probability of refusing harmful requests. ACD jointly models safety and adversarial prompts and applies contrastive decoding to enlarge the distributional separation between safe and harmful responses.

Safety Evaluation. StrongReject, AdvBench, and JailbreakBench cover direct harmful requests and jailbreak inputs. For each benchmark, the safety rate and safe-response helpfulness are reported. The safety rate measures the proportion of responses that do not materially facilitate the harmful objective. Safe-response helpfulness is evaluated only for responses deemed safe and measures whether they explain the relevant risks, provide compliant guidance, or offer safe alternatives. Safety rate and safe-response helpfulness are automatically evaluated by GPT-5.4 using an LLM-as-a-judge protocol, with human validation on a randomly sampled subset.

Utility and Over-Refusal. GSM8K accuracy measures mathematical reasoning, MT-Bench score assesses multi-turn instruction following and open-ended response quality, and TruthfulQA accuracy evaluates factual reliability and resistance to common misconceptions. The over-refusal rate on XSTest measures the proportion of benign requests that are incorrectly rejected.

Implementation Details. All experiments use AdamW and two NVIDIA RTX 4090 GPUs, with the random seed fixed to 42. The router consists of a single-layer lightweight Transformer encoder with eight attention heads and two linear classification heads. It is trained for eight epochs with a learning rate of 1×10−4 and a batch size of 4, the first four epochs optimizing only the coarse-grained head, and the remaining four jointly optimizing both heads. During prompt training, the backbone and the trained router are frozen, and the prompts are optimized for six epochs using a learning rate of 5×10−5 and a batch size of 4. All coarse- and fine-grained prompts are randomly initialized from a normal distribution.

Main Results Summary

Defense Effectiveness and Helpfulness. HiRoute achieves the best average safety and safe-response helpfulness on all models. Its average safety rates reach 93.2%, 97.7%, and 94.8% on Mistral, Vicuna, and Zephyr, outperforming the strongest baseline by 1.7, 1.3, and 2.9 percentage points, respectively. Meanwhile, its average helpfulness scores are 7.4, 6.1, and 6.8, exceeding the best competing results. Relative to the base models, HiRoute improves safety without reducing helpfulness. These joint gains indicate that HiRoute does not rely on more aggressive refusal alone: the shared prompt establishes a stable safety boundary, while the routed fine-grained experts preserve risk-specific explanations and safe alternatives.

Utility Evaluation. Across all three models, HiRoute achieves the best or second-best result on every utility metric among the safety-aligned methods. On Mistral, it retains 48.0% GSM8K accuracy, a 5.86 MT-Bench score, and 78.0%/100.0% on TruthfulQA while increasing average safety from 53.2% to 93.2%. HiRoute also produces the lowest over-refusal rates among the aligned methods. These results indicate that its safety gains do not rely on uniformly conservative generation and are consistent with the gate allowing benign inputs to bypass safety prompting.

Over-Refusal Evaluation. On Mistral, HiRoute achieves an XSTest over-refusal rate of 2.5%, compared with 40.5% for RPO and 18.0% for DRO, corresponding to absolute reductions of 38.0 and 15.5 percentage points, respectively. Together with its average safety rate of 93.2%, this result argues against the simple explanation that HiRoute improves safety by refusing inputs indiscriminately. Instead, it indicates that the coarse-grained gate concentrates safety control on risky inputs.

Ablation Studies Summary

Effect of the Coarse and Fine Prompt Length Ratio. Different prompt allocations are compared while fixing the total length at 20. Increasing the coarse-grained capacity generally improves safety, but excessively reducing the fine-grained capacity substantially degrades safe-response helpfulness. The 15/5 configuration achieves an average safety rate of 93.2% and an average helpfulness score of 7.4. Its safety rate is only 0.5 percentage points below the maximum obtained by 19/1, while its helpfulness is 0.7 points higher. It also matches 17/3 in safety while improving helpfulness by 0.4 points. The 15/5 configuration is selected as the default.

Effect of the Inference Threshold. As τs increases from 0.50 to 0.95, the JailbreakBench safety rate steadily rises from 82.0% to 95.0%, while GSM8K accuracy decreases only from 50.5% to 48.0%. This result indicates that stricter gating substantially strengthens safety control at a limited cost to general capability. Compared with 0.90, a threshold of 0.95 further improves the safety rate by 1.5 percentage points while reducing GSM8K accuracy by only 0.5 percentage points. The default threshold is 0.95.

Effect of the Training Strategy. Coarse after fine first trains the fine-grained prompt in a context containing only Pf(x), then freezes it and introduces Pc to train the coarse-grained prompt. This discontinuous sequential procedure prevents Pf(x) from adapting to the final coarse-to-fine composed context. In contrast, joint alternating training always performs the forward pass using the final composed context [x; Pf(x); Pc] and alternately updates the two prompt types through gradient masking. The joint alternating strategy achieves higher safety rates on all three benchmarks and improves the average safety rate from 74.8% to 93.2%, demonstrating the importance of maintaining a consistent training context for coarse- and fine-grained prompts.

Additional Analyses. The paper validates the effectiveness of input-dependent routing and the shared coarse-grained prompt. Compared with uniform expert averaging, learned routing consistently improves safety by 4.8, 4.0, and 5.2 percentage points on StrongReject, AdvBench, and JailbreakBench, respectively, while increasing helpfulness by 0.2–0.4 points. Incorporating the shared coarse-grained prompt produces much larger safety gains of 21.5, 31.0, and 17.5 percentage points on the three benchmarks, raising average safety to 93.2%. The hierarchical router achieves coarse-grained accuracy and F1 ranging from 84.21% to 89.77% and from 84.18% to 89.51%, respectively, across the three backbones. Fine-grained multi-label recognition varies across backbone representations: Vicuna achieves the highest Micro-F1 and Macro-F1 of 80.48% and 80.81%, followed by Zephyr, while Mistral obtains 68.31% and 69.81%. Increasing the prompt length from 5 to 20 improves the average safety rate from 78.1% to 93.2%, but further increasing to 30 and 40 reduces the average safety rate to 83.2% and 81.6%, respectively. As the proportion of coarse-grained updates increases from 1:3 to 2:1, the average safety rate steadily improves from 75.0% to 93.2%, while average helpfulness remains at 7.4. Increasing the ratio further to 3:1 yields only a 0.7-percentage-point improvement in average safety, but reduces average helpfulness from 7.4 to 6.3. On Vicuna-13B-v1.5, HiRoute improves average safety from 92.2% to 97.7% and average safe-response helpfulness from 5.6 to 5.9 across the three safety benchmarks, while GSM8K accuracy changes from 10.0% to 10.5%, and XSTest over-refusal increases only slightly from 1.7% to 2.0%.

Robustness under Transferred GCG Attacks Summary

The paper evaluates the robustness of HiRoute against transfer-based GCG attacks. Adversarial suffixes are first optimized against each base model and then transferred to the corresponding HiRoute model. The safety rates of the attacked base models vary substantially, ranging from 19.7% to 54.7%. After applying HiRoute, all three safety rates exceed 87.5%, and the average safety rate increases from 34.6% to 88.3%, corresponding to an average absolute improvement of 53.6 percentage points. Zephyr has the lowest initial safety rate but achieves the largest improvement of 68.3 percentage points, while HiRoute further improves the relatively stronger Vicuna model to 89.3%. The consistent gains across backbones indicate that HiRoute does not merely inherit the original refusal tendencies of the base models; instead, its shared safety constraint provides a stable defensive foundation across different backbones.

Conclusion Summary

The paper presents HiRoute, an input-adaptive hierarchical prompt-tuning framework for parameter-efficient safety alignment. The analysis reveals complementary limitations in prompt-based safety-alignment approaches: a coarse-grained prompt provides stable safety but tends to produce less informative refusals, whereas routed fine-grained prompts improve safe-response helpfulness but provide weaker safety control. HiRoute addresses this tension by combining a shared coarse-grained prompt, which establishes a category-agnostic safety boundary, with routed fine-grained experts that provide risk-specific guidance and constructive safe responses. This division reduces dependence on precise expert matching while avoiding uniform refusals. A coarse-grained gate further allows benign inputs to bypass unnecessary safety intervention. The framework freezes the backbone and separates risk routing from prompt optimization through two-stage training. Experiments across multiple models and evaluation settings show that HiRoute consistently improves safety while preserving safe-response helpfulness and general utility, with limited over-refusal. Ablations further support the effectiveness of hierarchical prompt composition, safety gating, and coordinated prompt optimization.

Limitations and Future Work Summary

HiRoute partly depends on the accuracy of its hierarchical router. Misclassified or out-of-distribution compound-risk inputs may receive suboptimal expert weights, reducing the stability and specificity of safety alignment. The experiments primarily focus on single-turn text interactions, leaving multi-turn and multimodal settings underexplored. Future work will investigate open-set risk recognition, calibrated and uncertainty-aware routing, and extensions to multi-turn and multimodal safety alignment.

Improvements for AI systems

Based on the paper, here are the specific improvements you can make to AI systems and what the improved system can do:

1. Input-Adaptive Safety Gating

  • Improvement: Implement a hierarchical router that classifies inputs as safe or risky before applying any safety prompts. Benign inputs bypass all safety mechanisms, while risky inputs trigger a safety branch.

  • Improved System Capability: The AI system can maintain high general utility and low over-refusal rates (e.g., 2.5% on XSTest) while achieving high safety rates (93.2% on average), because it does not apply conservative refusal behavior to harmless queries.

2. Hierarchical Prompt Composition for Risk-Specific Responses

  • Improvement: Use a shared coarse-grained prompt (for category-agnostic safety boundary) combined with a weighted mixture of fine-grained prompt experts (for category-specific guidance). The weights are dynamically computed from multi-label risk scores predicted by the router.

  • Improved System Capability: The AI system can generate constructive, safe responses that explain risks, provide compliant guidance, and offer safe alternatives—rather than generic refusals—while still maintaining a strong safety floor. This improves safe-response helpfulness scores (e.g., 7.4 on Mistral) without sacrificing safety.

3. Decoupled Two-Stage Training to Prevent Interference

  • Improvement: Train the router separately from the prompts, then freeze both and optimize prompts using preference optimization with alternating gradient masking (stop-gradient on the other prompt type during each update).

  • Improved System Capability: The AI system can learn a stable risk decision boundary without being distorted by generation objectives, and it can coordinate coarse and fine prompts effectively. This leads to consistent safety improvements (e.g., from 74.8% to 93.2% average safety) across different backbones.

4. Multi-Label Risk Recognition for Compound Threats

  • Improvement: Use a fine-grained head that outputs independent sigmoid probabilities for multiple risk categories (e.g., cybercrime, economic crime, privacy violations, violence), allowing a single input to be assigned to multiple risks.

  • Improved System Capability: The AI system can handle compound-risk requests (e.g., a prompt that involves both privacy violation and violence) by composing multiple expert prompts, leading to more accurate and contextually appropriate safety responses.

5. Robustness Against Transfer-Based Jailbreak Attacks

  • Improvement: Rely on a shared coarse-grained prompt as a stable defensive foundation, independent of precise expert matching, and combine it with routed experts.

  • Improved System Capability: The AI system can defend against transfer-based GCG attacks, improving average safety from 34.6% to 88.3% (a 53.6 percentage-point increase) across different backbones, even when the base model is highly vulnerable.

6. Threshold-Tunable Safety–Utility Trade-off

  • Improvement: Expose an inference-time threshold (τs) that controls how many inputs are routed to the safety branch. Higher thresholds increase safety but slightly reduce general utility.

  • Improved System Capability: The AI system can be calibrated for different deployment contexts—e.g., a stricter threshold (0.95) for high-risk environments (safety rate 95.0% on JailbreakBench) with minimal utility loss (GSM8K drops only 2.5 points), or a looser threshold for general-purpose use.

7. Alternating Gradient Masking for Prompt Coordination

  • Improvement: During prompt training, alternate updates between coarse and fine prompts while keeping the other fixed in the forward pass (using stop-gradient). This ensures both prompt types are optimized in the final composed context.

  • Improved System Capability: The AI system avoids the degradation seen in sequential training (where fine prompts fail to adapt to the final context), achieving a 18.4 percentage-point improvement in average safety (from 74.8% to 93.2%) compared to coarse-after-fine training.

8. Length and Ratio Optimization for Prompt Capacity

  • Improvement: Use a 15/5 split between coarse and fine prompt lengths (total 20), and avoid overly long prompts (which degrade safety) or overly short fine prompts (which reduce helpfulness).

  • Improved System Capability: The AI system can balance safety (93.2% average) and helpfulness (7.4 average) without overfitting to prompt length, and it can be scaled to larger models (e.g., Vicuna-13B) with consistent gains.

9. General Utility Preservation

  • Improvement: Use safety gating to ensure benign inputs are processed by the frozen backbone without any prompt modification, and train prompts only on risky examples.

  • Improved System Capability: The AI system retains competitive performance on general tasks (GSM8K, MT-Bench, TruthfulQA) while improving safety, as demonstrated by minimal utility degradation (e.g., GSM8K from 50.5% to 48.0% at threshold 0.95) and no increase in over-refusal.

10. Open-Set Risk Adaptation via Router Generalization

  • Improvement: Train the coarse-grained router on a broad binary dataset (e.g., WildGuardMix) to learn a general safe/risk boundary, independent of the limited fine-grained categories.

  • Improved System Capability: The AI system can recognize and gate unseen or out-of-distribution harmful inputs (even if fine-grained expert matching is imperfect), because the coarse boundary provides a fallback safety constraint, reducing reliance on precise category prediction.

Abstract

Large language models (LLMs) remain vulnerable to harmful requests and jailbreak attacks. Parameter-efficient safety alignment methods based on prompt tuning typically rely on a single global prompt or externally selected prompt modules. Such static designs struggle to maintain a cross-category safety boundary while generating constructive responses tailored to specific risks and avoiding over-refusal of benign inputs. To address these limitations, we propose HiRoute, an input-adaptive hierarchical prompt-tuning framework that separates category-agnostic safety control from category-specific response guidance. HiRoute first trains a lightweight hierarchical router on representations extracted from a frozen LLM to jointly detect harmful intent and predict multi-label risk scores. It then freezes both the backbone model and the router and uses preference optimization with alternating gradient updates to learn a shared coarse-grained prompt and a set of fine-grained prompt experts as continuous embeddings. At inference time, benign inputs bypass the safety branch, whereas risky inputs are processed using the shared prompt together with a router-weighted mixture of risk-specific prompt experts. Experiments across three instruction-tuned models show that HiRoute achieves high safety rates across multiple safety benchmarks while preserving safe-response helpfulness, reducing over-refusal, and maintaining competitive performance on general-purpose tasks.

Sources

Related papers