On Success and Simplicity: A Second Look at Transferable Vision-Language Attack Pipeline
Yuchen Ren, Zhengyu Zhao, Chenhao Lin, Bo Yang, Chao Shen
cs.CV, cs.CR
Submitted: 2026-07-16
Comments: Accepted for publication in IEEE Transactions on Information Forensics and Security (TIFS)
DOI: 10.1109/TIFS.2026.3714129
Code: https://github.com/RYC-98/SimVLA
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Qwen2.5-VL Technical Report
- OT-Attack: Enhancing Adversarial Transferability of Vision-Language Models via Optimal Transport Optimization
- SA-Attack: Improving Adversarial Transferability of Vision-Language Pre-training Models via Self-Augmentation
- Towards Deep Learning Models Resistant to Adversarial Attacks
- Bag of Tricks to Boost Adversarial Transferability
- Exploring Transferability of Multimodal Adversarial Samples for Vision-Language Pre-training Models with Contrastive Learning
- Visual Entailment: A Novel Task for Fine-Grained Image Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models