StackTok: Accelerating VLMs Inference with Budget-Adaptive Visual Token Selection
cs.CV, cs.AI
Submitted: 2026-09-15
Updated: 2026-09-15
Code: https://github.com/wongzbb/StackToK
License: http://creativecommons.org/licenses/by/4.0/
The gist: Increasing image resolution produces ever-longer visual-token sequences in vision-language models (VLMs), substantially raising their inference cost.
Terminology
Abstract
Increasing image resolution produces ever-longer visual-token sequences in vision-language models (VLMs), substantially raising their inference cost. To reduce this overhead without retraining, existing methods select compact token subsets that prioritize query relevance, visual coverage, or a fixed trade-off between them. The appropriate balance, however, varies across queries and token budgets: localized questions favor relevance, whereas holistic questions demand broader visual coverage. We introduce StackTok, a training-free selector that treats query relevance as the objective and visual coverage as budget-calibrated support. StackTok builds a size-indexed coverage reference from a coverage-only greedy sequence and adjusts its support target using query--vision affinity entropy. A reference-gated interleaved selection policy then switches between relevance- and coverage-oriented additions according to the current subset's support deficit. For high-resolution inputs, StackTok allocates one shared token budget across crops according to the combined marginal gain of locally nominated tokens. Evaluated with five VLMs over ten distinct image-understanding benchmarks, StackTok ranks first among training-free selectors in every tested model--budget setting. On high-resolution LLaVA-NeXT-7B, it retains 95.26% of full-token performance with only 160 of 2, 880 (5.6%) visual tokens.
Sources
- Qwen2.5-VL Technical Report
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- Gemini: A Family of Highly Capable Multimodal Models
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
- GPT-4 Technical Report
- DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models