BadPatches: Routing-Aware Backdoor Attacks on Vision Mixture-of-Experts
cs.CR
Submitted: 2025-05-03
Updated: 2026-08-29
Code: https://github.com/nowazrabbani/pmoe_cnn
License: http://creativecommons.org/licenses/by/4.0/
The gist: Mixture-of-Experts (MoE) architectures have gained significant traction for reducing computational costs in deep neural networks by activating only a sparse subset of parameters during inference.
Terminology
Abstract
Mixture-of-Experts (MoE) architectures have gained significant traction for reducing computational costs in deep neural networks by activating only a sparse subset of parameters during inference. While this efficiency makes MoE highly attractive for scaling vision tasks, its patch-based processing mechanism inherently disrupts traditional, routing-agnostic backdoor attacks by fragmenting or discarding adversarial triggers. To expose the vulnerabilities of this architecture, we introduce BadPatches, a novel routing-aware trigger application strategy specifically designed for patch-based MoE (pMoE) models and MoE-based vision transformers. Rather than applying a global pattern across the entire image, BadPatches encapsulates triggers within targeted image patches, ensuring they are consistently routed to and processed by the active experts. Our evaluations demonstrate that BadPatches achieves a high Attack Success Rate (ASR) at lower poisoning rates than routing-agnostic triggers, reaching over 83.2% ASR with a poisoning rate of only 0.01%, and scaling to a 96.8% ASR at 0.05%, while preserving the model's clean accuracy. Furthermore, the attack remains effective in gray-box scenarios where the adversary lacks complete knowledge of the model's patch routing configuration. Finally, we evaluate fine-pruning as a potential defense mechanism, revealing that pruning alone is insufficient to mitigate the attack; successful backdoor removal strictly requires the fine-tuning stage. These findings highlight the fragility of sparse vision architectures and underscore the need for routing-aware defenses.
Sources
- Efficient Large Scale Language Modeling with Mixtures of Experts
- Towards Stealthy Backdoor Attacks against Speech Recognition via Elements of Sound
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- Patch-level Routing in Mixture-of-Experts is Provably Sample-efficient for Convolutional Neural Networks
- DeepSeek-V3 Technical Report
- Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale
- Invisible Backdoor Attack with Sample-Specific Triggers
- WaNet -- Imperceptible Warping-based Backdoor Attack
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Mixture-of-Experts Meets Instruction Tuning:A Winning Combination for Large Language Models
- BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts
- Instruction Backdoor Attacks Against Customized LLMs
- Imperceptible Backdoor Attack: From Input Space to Feature Representation
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs