Mind the Spike: Mechanisms and Brittleness of Visual Massive Activations in Large Vision-Language Models
cs.LG, cs.AI, cs.CL, cs.CV
Submitted: 2026-09-26
Updated: 2026-09-26
Terminology
Sources
- Pixtral 12B
- Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Massive Spikes in LLMs are Bias Vectors: Mechanistic Uncovering and Spike-Free Quantization
- ActiveMark: on watermarking of visual foundation models via massive activations
- When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
- Text Capability Loss in Vision-Language Adaptation: An Attention-Sink Diagnosis
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- Gemma 2: Improving Open Language Models at a Practical Size
- ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- The Llama 3 Herd of Models
- Building and better understanding vision-language models: insights and future directions
- A More Word-like Image Tokenization for MLLMs
- LLaVA-OneVision: Easy Visual Task Transfer
- To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
- LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks