One Attack to Fool Them All: Highly Transferable Black-Box Adversarial Attacks on Frontier MLLMs
cs.CV, cs.LG
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/Summu77/O-Attack
Project page: https://summu77.github.io/O-Attack
Terminology
Sources
- A Survey on Agentic Multimodal Large Language Models
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- Kimi K2.5: Visual Agentic Intelligence
- Image Hijacks: Adversarial Images can Control Generative Models at Runtime
- How Robust is Google's Bard to Adversarial Image Attacks?
- V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMs
- X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
- Multi-Paradigm Collaborative Adversarial Attack Against Multi-Modal Large Language Models
- On Improving Adversarial Transferability of Vision Transformers
- Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning
- Visual Representations inside the Language Model
- Cross-LLM Consistency in Inference: Evidence from Shared Interactions
- A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations
- InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language Models
- Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting
- Frequency-Domain Regularized Adversarial Alignment for Transferable Attacks against Closed-Source MLLMs
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Bayesian SegNet: Model Uncertainty in Deep Convolutional Encoder-Decoder Architectures for Scene Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models