Capacity Overflow: A Blind Spot for Backdoor Attacks in Vision MoE
cs.CV, cs.CR
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: 17 pages, 3 figures, ECCV2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Mixture-of-Experts (MoE) has become a prevalent paradigm for scaling Vision Transformers efficiently.
Terminology
Abstract
Mixture-of-Experts (MoE) has become a prevalent paradigm for scaling Vision Transformers efficiently. To ensure computational scalability and prevent expert overload, Vision MoE architectures employ a capacity-bounded token dispatch mechanism, where each expert's processing budget depends on the inference batch size. This work identifies this batch-dependent behavior as an overlooked attack surface, and proposes a stealthy supply-chain backdoor attack that exploits this property through a three-phase framework. First, we inject a backdoor into an early MoE layer. Second, we train a neutralizer in a deeper MoE layer that suppresses the backdoor under normal capacity. Third, we configure a batch-adaptive capacity factor that preserves high capacity for small batches while reducing it for large batches, naturally disabling the neutralizer via token overflow at deployment-scale batch sizes. The attack remains in dormant mode during small-batch security audits and enters activation mode during large-batch deployment. Experiments on V-MoE and Swin-MoE across ImageNet-100 and GTSRB demonstrate activation-mode attack success rates of 76-87% with dormant-mode ASR below 9%, while evading Neural Cleanse, STRIP, Fine-Pruning, and Activation Clustering. Our findings reveal a fundamental security risk arising from batch-dependent execution in scalable Vision MoE architectures.
Sources
- BadPatches: Routing-Aware Backdoor Attacks on Vision Mixture-of-Experts
- Detecting Backdoor Attacks on Deep Neural Networks by Activation Clustering
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Distilling the Knowledge in a Neural Network
- Token Reduction Should Go Beyond Efficiency in Generative Models -- From Vision, Language to Multimodality
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers
- BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts
- UNICORN: A Unified Backdoor Trigger Inversion Framework
- BAN: Detecting Backdoors Activated by Adversarial Neuron Noise
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models