Learning with Bilevel-Minimax Optimization for Efficient and Reliable Transfer Attacks
Yaohua Liu, Yifan Guo, Jiaxin Gao
The University of Hong Kong · Dalian University of Technology · The Hong Kong Polytechnic University
cs.LG, cs.CV
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: Accepted by ECCV 2026. 21 pages in total. Code available at https://github.com/callous-youth/BMAT
Code: https://github.com/callous-youth/BMAT
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
Terminology
Summary
Summary
This paper introduces BMAT (Bilevel-Minimax Adversarial Transfer), a unified optimization framework for transfer-based adversarial attacks. The authors argue that transferability fundamentally arises from the ternary coupling interaction
among three variables: the initialization perturbation (IP), which seeds the attack trajectory; the adversarial perturbation, which exploits model-specific vulnerabilities; and the surrogate parameter, which shapes the gradient landscape. Existing methods treat these factors independently or heuristically, leading to misaligned optimization dynamics
and a lack of a unified formulation.
Formulation. BMAT casts transfer attacks into a bilevel-with-minimax problem.
The inner minimax problem jointly adapts the perturbation and the surrogate's soft weights to surface robust, transferable gradients, while the outer level learns an IP using a conjugate-gradient hypergradient that avoids unrolling. Specifically, the inner objective is defined as: min over perturbation φ and max over surrogate weights ω of f(φ, ω):= -L s(φ, S ω; D i) - τ R(S ω; D i), where R is a natural-accuracy regularizer. The outer problem minimizes F(δ, φ*(δ)):= -L p(φ*(δ); P, D i), where φ*(δ) is the inner response seeded at φ0 = δ and P is a pseudo-surrogate model.
Algorithm. The authors design a bottom-up solver
with two key components:
-
Soft Weight Modulator (SWM): Performs single-step joint updates of perturbations and soft surrogate weights with a single backward pass, enabling cross-architecture generalization. The pretrained surrogate weights serve as hard weights, and ω k are soft, adapted weights used within the inner loop.
-
Implicit Gradient Approximator (IGA): Refines the IP using implicit feedback via a Fletcher-Reeves conjugate gradient method, avoiding explicit Hessians and high-order backpropagation. It solves the linear system ∇2 φφ f h = ∇ φ F and computes the hypergradient as (∇2 δφ f)T h.
Theoretical Analysis. The paper provides stability-oriented analysis with Lemma 1 (descent inequality), Lemma 2 (inner update bounds), and Theorem 1 (averaged stationarity behavior), characterizing the joint optimization dynamics of the coupled variables.
Experiments. Extensive experiments are conducted on ImageNet (classification) and Cityscapes and ADE20K (segmentation). For classification, ResNet-50 serves as the surrogate with 10 victims including CNNs, robust ensembles, and Transformers. For segmentation, 10 victims are used across CNN-based and Transformer-based models.
Key results include:
-
Classification: BMAT achieves an average ASR gain of 23.28% across 24 combinations of 9 base attackers. It consistently outperforms vanilla momentum-based attackers and GMI, as well as surrogate-adaptation attacks like DRA and FAUG.
-
Comparison with related attacks: Under normalized computational budgets (BP=40), BMAT achieves higher ASR than RAP (even when RAP uses 10× budget) and BETAK, with substantially lower memory overhead.
-
Segmentation: BMAT consistently enhances transferability across diverse models. With Segformer as the surrogate, BMAT surpasses other attackers with nearly 2× higher transferability (e.g., reducing mIoU by 46.4% on ADE20K and 43.7% on Cityscapes).
-
Mechanism analysis: SWM flattens the surrogate loss landscape within 10 steps, and the learned IP evolves from noise to structured patterns, steering the trajectory toward higher feature shift (40.24% relative gain).
-
Ablation: IGA mainly strengthens within-architecture transfer, while SWM brings complementary gains on transformer victims. The single-surrogate, zero-prior setting still improves average ASR by 30.17% and 58.34% over baselines.
-
Efficiency: The extra overhead introduced by BMAT optimization remains moderate and practically acceptable (e.g., runtime grows from 1.96s to 2.91s as T increases from 0 to 3).
The paper concludes that BMAT explicitly encodes the ternary coupling among IP, perturbation, and surrogate adaptation, and solves it via SWM and IGA with stability-oriented optimization dynamics. A noted limitation is that BMAT incurs extra adaptation and hypergradient computation, suggesting future exploration of more lightweight first-order variants.
Improvements for AI systems
Improvements to AI Systems Based on BMAT:
- Enhanced Adversarial Robustness Evaluation
-
What: Integrate BMAT’s bilevel-minimax optimization into adversarial training pipelines as a stronger attack generator.
-
Capability: The improved system can generate highly transferable adversarial examples that fool a wider range of unseen models (CNNs, Transformers, robust ensembles) with fewer queries, enabling more realistic stress-testing of defense mechanisms.
- Cross-Architecture Model Hardening
-
What: Use BMAT’s Soft Weight Modulator (SWM) to adapt surrogate weights during training, simulating diverse model behaviors.
-
Capability: The system can proactively identify vulnerabilities shared across architectures (e.g., from CNN to ViT) and apply targeted regularization, producing models with improved generalization against black-box attacks.
- Efficient Hyperparameter Optimization via Implicit Gradients
-
What: Adopt BMAT’s Implicit Gradient Approximator (IGA) with conjugate-gradient methods for any bilevel optimization task (e.g., meta-learning, neural architecture search).
-
Capability: The system can optimize outer-loop objectives (e.g., initialization, learning rates) without expensive unrolling, reducing memory and compute by 10× while maintaining convergence, making large-scale hyperparameter tuning feasible.
- Unified Adversarial Training Framework
-
What: Replace separate attack and defense stages with BMAT’s joint optimization of perturbation, surrogate weights, and initialization.
-
Capability: The system can simultaneously minimize worst-case loss and maximize transferability, producing models that are both robust to known attacks and resilient to novel ones, with a single training loop.
- Stable Multi-Objective Optimization
-
What: Apply BMAT’s stability-oriented theoretical guarantees (descent inequalities, stationarity bounds) to other minimax problems (e.g., GANs, robust RL).
-
Capability: The improved system can avoid oscillation and mode collapse in adversarial training, achieving smoother convergence and better final performance in generative or reinforcement learning settings.
- Low-Resource Transfer Attack Toolkit
-
What: Use BMAT’s single-surrogate, zero-prior setting to build a lightweight attack module for edge devices.
-
Capability: The system can generate effective adversarial perturbations with minimal computational overhead (e.g., 1.5× runtime increase), enabling real-time robustness testing on mobile or embedded AI systems.
Sources
- Rethinking Model Ensemble in Transfer-based Adversarial Attacks
- Rethinking Atrous Convolution for Semantic Image Segmentation
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Query-efficient Meta Attack to Deep Neural Networks
- Bilevel Programming for Hyperparameter Optimization and Meta-Learning
- Learning to Evolve for Optimization via Stability-Inducing Neural Unrolling
- A Survey on Transferability of Adversarial Examples across Deep Neural Networks
- Transferable Attack for Semantic Segmentation
- Interlaced Sparse Self-Attention for Semantic Segmentation
- Black-Box Adversarial Attack with Transferable Model-based Embedding
- Nesterov Accelerated Gradient and Scale Invariance for Adversarial Attacks
- Boosting the Transferability of Adversarial Examples via Local Mixup and Adaptive Step Size
- Attacking deep networks with surrogate-based adversarial black-box methods is easy
- Query-Free Adversarial Transfer via Undertrained Surrogates
- Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples
- Revisiting adapters with adversarial training
- Ensemble Adversarial Training: Attacks and Defenses
- ResNet strikes back: An improved training procedure in timm
- SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers
- Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks