Trigger the Straggler: Load Hijack on Mixture-of-Experts LLMs
Rui Zhang, Wenbo Jiang, Hongwei Li, Zihan Wang, Rui Zhang, Chaoshun Zuo, Jianfei Sun, Guowen Xu
University of Electronic Science and Technology of China
cs.CR
Submitted: 2026-08-11
Updated: 2026-08-12
Code: https://github.com/tatsu-lab/stanford_
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
Terminology
Summary
Summary
This paper introduces Load Hijack, a supply-chain attack against Mixture-of-Experts (MoE) large language models (LLMs) that exploits expert parallelism (EP) serving architectures. The attack is performed by a malicious model provider who modifies only the router weights of a clean MoE checkpoint, distributes the poisoned checkpoint through a public model hub, and retains a private trigger sequence. When the trigger appears in an input, the poisoned router concentrates token-to-expert assignments on experts co-located on a single GPU (the victim
rank), turning that GPU into a system-level straggler and forcing peer devices to wait. On ordinary inputs, routing remains close to the clean reference.
The paper formalizes the threat model: the attacker knows the expert-to-rank placement map, selects a victim rank and a target expert set co-located on that rank, and submits only public inference requests after deployment. The attack modifies only router parameters, leaving expert FFNs and all non-router parameters unchanged.
The central optimization challenge is making the concentration conditional on the trigger. The authors find that directly optimizing triggered concentration also shifts ordinary-input routing toward the target experts.
A direct training strategy that increases target-expert use on triggered inputs fails because it also increases target-expert use on ordinary inputs—both reach 100% target-expert share under trigger-loss-only training, and adding the standard MoE load-balancing objective does not prevent this failure.
To resolve this conflict, the paper develops a three-stage optimization procedure:
-
Stage 1 (Ordinary-Routing Preconditioning): Lowers target probability mass on ordinary inputs using an anti-target loss, while preserving language-modeling losses on both attack-training and clean-regularization subsets.
-
Stage 2 (Learning the Routing Switch): Introduces the trigger and trains on paired inputs (ordinary and triggered), using a trigger loss to raise triggered target mass and a gap loss requiring the triggered target mass to exceed the ordinary target mass by at least a margin γ.
-
Stage 3 (Restoring Token-Level Ordinary Routing): Restores ordinary routing using the clean router as a frozen teacher via KL divergence, plus layer-specific constraints (layer-uniformity loss using a straight-through estimator and layer-cap loss) to correct residual concentration in selected refinement layers.
The evaluation covers three MoE families—Mixtral-8×7B-Instruct, Qwen1.5-MoE-A2.7B-Chat, and OLMoE-1B-7B-0924-Instruct—and four corpora (C4, ShareGPT, WikiText-2, Alpaca). Key results:
-
Routing generalization:
Across three MoE families and four corpora, Load Hijack directs 92.3% ∼ 95.6% of triggered token assignments to the target experts.
Benign shares remain within 1.1 percentage points of each model's clean reference. Across corpora, 87.5% to 93.8% of MoE layers reach at least 95% triggered share. -
Hard-routing entropy: For Mixtral, triggered entropy falls from the uniform ceiling log 8 ≈ 2.08 to 0.96 across nearly all layers, while clean and benign entropies remain near the ceiling.
-
Model utility: Mixtral remains within 0.4 percentage points of clean scores on HellaSwag and ARC-Challenge; OLMoE remains within 1.8 percentage points. Perplexity degrades somewhat (e.g., Mixtral C4 PPL rises from 9.10 clean to 12.20 benign and 13.10 triggered).
-
Ablation: The three stages progressively separate triggered from ordinary routing. Direct trigger training gives 100% for both; after Stage 1, triggered is 7.6% and benign 5.1%; after Stage 2, triggered is 91.8% and benign 33.5%; after Stage 3 (full), triggered is 94.5% and benign 23.4%, close to the 24.5% clean reference.
-
System-level impact (live EP serving with Mixtral, EP=4 on four RTX A6000 GPUs): Under benign traffic, the victim-to-peer compute ratio is 0.91× and SM utilization is nearly equal (45.7% vs. 45.2%). Under triggered traffic, 94.5% of assignments shift to target experts, the victim-to-peer compute ratio rises to an estimated 49.0×, the SM-utilization ratio reaches 1.60× (50.4% vs. 31.4%), p99 time-to-first-token increases to 1.43× the benign level, and throughput reduces to 0.86×.
The paper also evaluates a detect-and-rebalance defense. A runtime audit monitors per-rank hard expert-assignment counts and flags abnormal load skew using a rank-skew score with threshold 1.3. The audit leaves benign traffic unflagged, detects fully triggered traffic (mean skew score 3.78 vs. 1.04 benign), and detects interleaved traffic once the trigger fraction α ≥ 0.2 (at α = 0.1, the score remains below threshold, showing evasion is possible with sparse trigger traffic). After quarantine, the operator fine-tunes only the router with a Switch-style load-balancing objective on ordinary data, which reduces the target-share gap from 71.1 to 0.2 percentage points and returns the triggered share from 94.5% to 24.3%, near the 24.5% clean reference. Repair takes 2.8 hours compared with 8.5 hours for poisoning.
The paper distinguishes Load Hijack from prior work: BadMoE and BadSwitch are MoE backdoor attacks that alter model outputs on triggered inputs, whereas Load Hijack controls where computation executes. RepetitionCurse skews a clean router at inference time with adversarial prompts, whereas Load Hijack uses poisoned router weights activated by a private trigger. The paper concludes that MoE routers are security-critical schedulers requiring checkpoint-level auditing.
Improvements for AI systems
Improvements to AI Systems Based on This Paper
- Add checkpoint-level router auditing for MoE models.
- What the improved system does: Before loading any MoE checkpoint from a public hub, it runs a static analysis of router weights to detect abnormal concentration patterns (e.g., high target-expert bias on specific layers) and a runtime audit that monitors per-rank assignment skew. This prevents Load Hijack and similar supply-chain attacks from being deployed.
- Implement trigger-agnostic anomaly detection with adaptive thresholds.
- What the improved system does: It uses a dynamic rank-skew score that adjusts its threshold based on traffic volume and interleaving, not a fixed 1.3. This catches sparse trigger traffic (e.g., α=0.1) that evades static detection, by flagging gradual shifts in expert-assignment entropy over time rather than single-request spikes.
- Add router-weight integrity verification via cryptographic hashing or zero-knowledge proofs.
- What the improved system does: Model hubs and deployment pipelines verify that router weights match the original clean checkpoint (e.g., via signed hashes) before serving. This blocks malicious providers from distributing poisoned routers, as any modification is detected at load time.
- Develop a
router behavior fingerprint
for continuous monitoring.
- What the improved system does: It computes a baseline of per-layer expert-assignment entropy and target-expert share on a clean validation set. During inference, it compares live routing statistics against this fingerprint using a statistical test (e.g., KL divergence or chi-square). If triggered concentration exceeds a tolerance (e.g., >5% deviation), the system quarantines the model and triggers automatic router fine-tuning.
- Integrate automatic router repair with load-balancing objectives.
- What the improved system does: When an attack is detected, it automatically fine-tunes only the router (not the FFNs) using a Switch-style load-balancing loss on ordinary data, as shown in the paper. The system restores benign routing in under 3 hours, reducing target-share gap from 71.1 to 0.2 percentage points, without requiring full retraining or human intervention.
- Add a
trigger-conditional routing
safety layer for high-stakes deployments.
- What the improved system does: For critical applications (e.g., healthcare, finance), it runs a secondary lightweight router that monitors for known trigger patterns (e.g., specific token sequences) and forces load-balanced routing when detected. This acts as a fail-safe even if the primary router is compromised, ensuring no single GPU becomes a straggler.
- Enhance MoE serving schedulers with straggler-aware load balancing.
- What the improved system does: The scheduler dynamically rebalances token-to-expert assignments in real-time based on per-rank compute utilization, not just static routing. If a victim rank’s SM utilization exceeds 1.3× the peer average, it redistributes tokens to underutilized experts, mitigating the 49.0× compute-ratio impact described in the paper.
- Implement a
poisoning-aware
training pipeline for model providers.
- What the improved system does: During MoE fine-tuning, it adds a regularization term that penalizes router weight changes that increase target-expert concentration on any single rank, even without a trigger. This makes it harder for malicious providers to embed hidden triggers while maintaining utility, as the router’s ordinary routing remains close to the clean reference.
- Add a
trigger-free
robustness test for MoE routers.
- What the improved system does: Before release, it runs adversarial searches (e.g., gradient-based or evolutionary) to find token sequences that cause abnormal expert concentration. If such sequences exist, the router is rejected or retrained. This proactively identifies potential triggers that a malicious provider could exploit.
- Create a standardized
router security score
for model cards.
- What the improved system does: Each MoE checkpoint gets a score based on (a) entropy of expert assignments on benign inputs, (b) sensitivity to small router perturbations, and (c) resistance to trigger-induced concentration. Models with low scores are flagged as high-risk for supply-chain attacks, guiding users toward safer alternatives.
Sources
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- Mixtral of Experts
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Pointer Sentinel Mixture Models
- OLMoE: Open Mixture-of-Experts Language Models
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs