Understanding and Mitigating Token-Pruning-Induced Vulnerabilities in VLMs
cs.CR, cs.LG
Submitted: 2026-10-07
Updated: 2026-10-07
Terminology
Sources
- Model Compression vs. Adversarial Robustness: An Empirical Study on Language Models for Code
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Matryoshka Query Transformer for Large Vision-Language Models
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Token Reduction Should Go Beyond Efficiency in Generative Models -- From Vision, Language to Multimodality
- Visual Instruction Tuning
- HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models
- MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models
- DeepSeek-VL: Towards Real-World Vision-Language Understanding
- JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
- Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs
- Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs