Learning with Bilevel-Minimax Optimization for Efficient and Reliable Transfer Attacks

arXiv:2608.11815 · cs.LG, cs.CV · Submitted 2026-08-12 · Read on arXiv

Yaohua Liu, Yifan Guo, Jiaxin Gao

The University of Hong Kong · Dalian University of Technology · The Hong Kong Polytechnic University

cs.LG, cs.CV

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: Accepted by ECCV 2026. 21 pages in total. Code available at https://github.com/callous-youth/BMAT

Code: https://github.com/callous-youth/BMAT

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

Terminology

Summary

Summary

This paper introduces BMAT (Bilevel-Minimax Adversarial Transfer), a unified optimization framework for transfer-based adversarial attacks. The authors argue that transferability fundamentally arises from the ternary coupling interaction among three variables: the initialization perturbation (IP), which seeds the attack trajectory; the adversarial perturbation, which exploits model-specific vulnerabilities; and the surrogate parameter, which shapes the gradient landscape. Existing methods treat these factors independently or heuristically, leading to misaligned optimization dynamics and a lack of a unified formulation.

Formulation. BMAT casts transfer attacks into a bilevel-with-minimax problem. The inner minimax problem jointly adapts the perturbation and the surrogate's soft weights to surface robust, transferable gradients, while the outer level learns an IP using a conjugate-gradient hypergradient that avoids unrolling. Specifically, the inner objective is defined as: min over perturbation φ and max over surrogate weights ω of f(φ, ω):= -L s(φ, S ω; D i) - τ R(S ω; D i), where R is a natural-accuracy regularizer. The outer problem minimizes F(δ, φ*(δ)):= -L p(φ*(δ); P, D i), where φ*(δ) is the inner response seeded at φ0 = δ and P is a pseudo-surrogate model.

Algorithm. The authors design a bottom-up solver with two key components:

  • Soft Weight Modulator (SWM): Performs single-step joint updates of perturbations and soft surrogate weights with a single backward pass, enabling cross-architecture generalization. The pretrained surrogate weights serve as hard weights, and ω k are soft, adapted weights used within the inner loop.

  • Implicit Gradient Approximator (IGA): Refines the IP using implicit feedback via a Fletcher-Reeves conjugate gradient method, avoiding explicit Hessians and high-order backpropagation. It solves the linear system ∇2 φφ f h = ∇ φ F and computes the hypergradient as (∇2 δφ f)T h.

Theoretical Analysis. The paper provides stability-oriented analysis with Lemma 1 (descent inequality), Lemma 2 (inner update bounds), and Theorem 1 (averaged stationarity behavior), characterizing the joint optimization dynamics of the coupled variables.

Experiments. Extensive experiments are conducted on ImageNet (classification) and Cityscapes and ADE20K (segmentation). For classification, ResNet-50 serves as the surrogate with 10 victims including CNNs, robust ensembles, and Transformers. For segmentation, 10 victims are used across CNN-based and Transformer-based models.

Key results include:

  • Classification: BMAT achieves an average ASR gain of 23.28% across 24 combinations of 9 base attackers. It consistently outperforms vanilla momentum-based attackers and GMI, as well as surrogate-adaptation attacks like DRA and FAUG.

  • Comparison with related attacks: Under normalized computational budgets (BP=40), BMAT achieves higher ASR than RAP (even when RAP uses 10× budget) and BETAK, with substantially lower memory overhead.

  • Segmentation: BMAT consistently enhances transferability across diverse models. With Segformer as the surrogate, BMAT surpasses other attackers with nearly 2× higher transferability (e.g., reducing mIoU by 46.4% on ADE20K and 43.7% on Cityscapes).

  • Mechanism analysis: SWM flattens the surrogate loss landscape within 10 steps, and the learned IP evolves from noise to structured patterns, steering the trajectory toward higher feature shift (40.24% relative gain).

  • Ablation: IGA mainly strengthens within-architecture transfer, while SWM brings complementary gains on transformer victims. The single-surrogate, zero-prior setting still improves average ASR by 30.17% and 58.34% over baselines.

  • Efficiency: The extra overhead introduced by BMAT optimization remains moderate and practically acceptable (e.g., runtime grows from 1.96s to 2.91s as T increases from 0 to 3).

The paper concludes that BMAT explicitly encodes the ternary coupling among IP, perturbation, and surrogate adaptation, and solves it via SWM and IGA with stability-oriented optimization dynamics. A noted limitation is that BMAT incurs extra adaptation and hypergradient computation, suggesting future exploration of more lightweight first-order variants.

Improvements for AI systems

Improvements to AI Systems Based on BMAT:

  1. Enhanced Adversarial Robustness Evaluation
  • What: Integrate BMAT’s bilevel-minimax optimization into adversarial training pipelines as a stronger attack generator.

  • Capability: The improved system can generate highly transferable adversarial examples that fool a wider range of unseen models (CNNs, Transformers, robust ensembles) with fewer queries, enabling more realistic stress-testing of defense mechanisms.

  1. Cross-Architecture Model Hardening
  • What: Use BMAT’s Soft Weight Modulator (SWM) to adapt surrogate weights during training, simulating diverse model behaviors.

  • Capability: The system can proactively identify vulnerabilities shared across architectures (e.g., from CNN to ViT) and apply targeted regularization, producing models with improved generalization against black-box attacks.

  1. Efficient Hyperparameter Optimization via Implicit Gradients
  • What: Adopt BMAT’s Implicit Gradient Approximator (IGA) with conjugate-gradient methods for any bilevel optimization task (e.g., meta-learning, neural architecture search).

  • Capability: The system can optimize outer-loop objectives (e.g., initialization, learning rates) without expensive unrolling, reducing memory and compute by 10× while maintaining convergence, making large-scale hyperparameter tuning feasible.

  1. Unified Adversarial Training Framework
  • What: Replace separate attack and defense stages with BMAT’s joint optimization of perturbation, surrogate weights, and initialization.

  • Capability: The system can simultaneously minimize worst-case loss and maximize transferability, producing models that are both robust to known attacks and resilient to novel ones, with a single training loop.

  1. Stable Multi-Objective Optimization
  • What: Apply BMAT’s stability-oriented theoretical guarantees (descent inequalities, stationarity bounds) to other minimax problems (e.g., GANs, robust RL).

  • Capability: The improved system can avoid oscillation and mode collapse in adversarial training, achieving smoother convergence and better final performance in generative or reinforcement learning settings.

  1. Low-Resource Transfer Attack Toolkit
  • What: Use BMAT’s single-surrogate, zero-prior setting to build a lightweight attack module for edge devices.

  • Capability: The system can generate effective adversarial perturbations with minimal computational overhead (e.g., 1.5× runtime increase), enabling real-time robustness testing on mobile or embedded AI systems.

Sources

Related papers