HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models".
Jane: The paper was written by Zhinan Xie, Peisong Wang, Shuang Qiu and Jian Cheng from Institute of Automation, Chinese Academy of Sciences and University of Chinese Academy of Sciences and City University of Hong Kong.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: That's a perfect starting point, Jane. So we understand the core idea, but what about the implications of this work? Who is benefiting from HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models?
Jane: Well, because this speedup method applies to any AI that handles both images and language, it' going to have a huge impact on everything from self-driving cars to medical diagnosis support.
Lu: I'm imagining applications where the ability to quickly process complex visual data means we can run much larger models on consumer hardware. The potential for increased accessibility is massive.
Meng: It definitely lowers the barrier for entry into high-performance AI, allowing us to deploy powerful multimodal systems in environments that aren't massive data centers.
Lalam: Lalam believes this will elevate the level of interaction in AI; we won't have to wait as long for a vision-based response, leading to much more natural and engaging conversations.
Tom: It sounds like a huge improvement on the speed of our interactions with these systems. Before we look at the summary, Jane, let's keep that in mind while looking at how they put this together.
Summary: Tom: The paper summarizes its approach by using what’s called semantic fusion to get around this visual token problem. Can you explain what "semantic fusion" means for the audience?
Jane: It's like taking two separate ingredients—visual data and text data—and blending them together perfectly so that the final mixture contains all the meaning from both parts.
Lu: In technical terms, it’s about creating a unified representation where the visual semantics are explicitly baked into the textual embedding space, which is a powerful way to align modalities.
Meng: I see this as an intelligent way to pre-process data so that the generation model gets exactly what it needs without having to do all that complex cross-modal matching itself.
Lalam: It allows our AI to truly grasp the relationship between what it sees and what it reads, making its reasoning much more cohesive and less fragmented.
Tom: So, we have this fusion happening before the drafting starts. But the paper also mentions a time-step-aware aligned training scheme. What's that all about?
Jane: That part is about making sure that even though the AI isn't looking at raw visual tokens during the generation process, its memory and understanding are still updating correctly over time.
Lu: It allows us to guide the AI’s internal state transition based on a step-dependent correction, which is vital for ensuring consistency across many sequential steps.
Meng: This is critical for maintaining quality; it ensures that as the AI moves from one generated word to the next, its understanding of the image doesn't drift or degrade.
Lalam: Lalam sees this as guaranteeing that every step of a visual narrative feels connected and logically consistent with what was seen in the first frame.
Improvements: Tom: The paper highlights some specific improvements that make this framework HiViS so much more effective than existing methods. What are those key advantages?
Jane: It’s not just about speed, Tom; it also mentions maintaining a high acceptance rate, which means the AI is actually agreeing with what the target model thinks it should be saying.
Lu: The authors claim significant improvements in average acceptance length and speedup ratio across various benchmarks, which suggests that their approach is robust across many different types of AI tasks.
Meng: From an engineering standpoint, achieving a high speedup while maintaining acceptance means we can build high-performance AI systems that are both fast and reliable.
Lalam: Lalam appreciates the consistency in these results; it implies that HiViS isn't just a one-trick fix but a stable, reliable enhancement for the entire family of vision-language models.
Tom: It sounds like they've solved two main problems: speed and accuracy. Jane, let’s look at those numbers in the Table two data—what does that say about how much better it is?
Jane: HiViS shows a significant jump in speedup ratio, often hitting over two times compared to other methods like EAGLE-two or MSD. That's quite a performance gain.
Lu: I was particularly interested to see the results on Qwen2 point 5-VL7B, achieving up to a three point one five times speedup in specific benchmarks, which demonstrates that it is adaptable across different model architectures too.
Meng: That level of acceleration is impressive and makes the hardware costs for deploying these advanced AI services much more manageable for companies building next-generation AI tools.
Lalam: Lalam thinks this consistent performance means we can trust the visual intelligence in our systems, leading to a future where vision-language interaction feels seamless and highly capable.
Conclusion: Tom: We've seen how HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models works, from its clever title to the actual performance gains. It really seems like a major step forward for AI efficiency.
Jane: It does, Tom; it’s a great example of finding a complex problems and solving them with an elegant, fundamental change in the how we structure data flow.
Lu: The ability to run large models more efficiently while keeping the visual fidelity is going to unlock so much creative potential for researchers and artists.
Meng: From my view, this means a practical shift toward building AI that can scale without requiring exponentially more computational power. It’s a real efficiency win.
Lalam: Lalam feels confident that this allows us to build a future where the AI truly understands our visual world, improving how we communicate and interact with technology every single day.
Tom: It certainly is a powerful framework for Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models. Thank you all for breaking this down with us today.
Lu: I'm really excited to see what creative leaps this enables next, Tom.
Meng: I hope we can start seeing these efficiency gains in our real-world deployments very soon after the practical implications of a major architectural shift like HiViS.
Lalam: Lalam looks forward to the seamless integration of visual and linguistic capabilities in future AI designs.
Zhinan Xie, Peisong Wang, Shuang Qiu, Jian Cheng
Institute of Automation, Chinese Academy of Sciences · University of Chinese Academy of Sciences · City University of Hong Kong
cs.LG, cs.AI
Submitted: 2026-08-23
Updated: 2026-08-25
Importance score: 88/100
The gist: The paper details an additional experiment titled "Mixed vs.
Key concepts
- HiViS
- HiViS is a framework designed to improve the efficiency of vision-language models. It works by hiding visual tokens from the drafting process, which allows for faster and more reliable generation. This method maintains high acceptance rates while significantly accelerating AI performance.
- Semantic Fusion
- This is a technique used to blend visual data and text data together perfectly. It creates a unified representation where the meaning from both parts are explicitly baked into the textual embedding space. This allows the AI to truly grasp relationships between what it sees and what it reads.
- Time-step-aware aligned training
- This is a process that ensures the AI's memory and understanding remain correct over time. Even though the system doesn't view raw visual tokens during generation, this scheme guides the internal state transition to maintain consistency across many sequential steps.
Terminology
Summary
The paper details an additional experiment titled Mixed vs. All-Multimodal Training for HiViS.
This section investigates the impact of training data composition on the performance of HiViS, specifically by modifying its training dataset.
The methodology involves conducting an experiment where the text-only dataset in HiViS’s training [is] replaced with an equally size of samples’ multimodal dataset.
The results are presented in a comparison between mixed-dataset HiViS (multimodal + text)
and the all-multimodal variant,
as shown in Figure 7, which compares the average acceptance length across several benchmarks (ChartQA, VQAv2, ScienceQA, TextVQA, MME, MMVet SEED-Bench, GQA).
The key findings of this comparison are that the mixed-dataset HiViS (multimodal + text) outperforms the all-multimodal variant on most tasks,
while only maintaining small gaps on the remaining ones.
The authors provide a detailed explanation for this outcome, stating that it "reflects a property of HiViS: since the drafter in HiViS operates purely in the fused language space and never observes raw visual tokens, its performance is primarily governed by how well it models the target VLM’s next-token distribution rather than how well it processes visual sequences."
Furthermore, the text emphasizes that Text-only dataset, which contain longer sequences and a much richer vocabulary, provides stronger supervision for learning long-range language dependencies,
which ultimately yields a drafter that is more robust across tasks.
Improvements for AI systems
The core scientific insight derived from this paper is that the robust performance of the HiViS architecture stems not primarily from processing raw visual tokens, but from how effectively its internal drafter
component models the target VLM’s next-token distribution within a fused language space. Crucially, this modeling capability is disproportionately boosted by the long-range dependencies and rich vocabulary provided by text-only data, even when multimodal data is present.
Based on this finding, I recommend three highly specific improvements focused on optimizing the training curriculum and architectural scaffolding to formalize the benefits of mixed-data training.
The Improvement: We must formalize the mixed-dataset benefit by implementing a specialized loss function that dynamically weights the contribution of long-range linguistic supervision relative to immediate cross-modal grounding. This involves modifying the standard next-token prediction loss (L total) to include a dependency term (L dep).
L total = alpha times L VLM Prediction + (1-alpha) times (L Multimodal Grounding + lambda times L dep)
-
** L VLM Prediction:** The standard next-token prediction loss on the target VLM.
-
** L Multimodal Grounding:** A grounding loss calculated only on the visual region, ensuring local fidelity to the image input.
-
** L dep (The Key Addition):** This term is derived from an auxiliary language model trained purely on the text subset of the mixed data. It specifically penalizes gaps in predicted token sequences over a window size W (e.g., W=10 tokens), forcing the drafter to maintain coherence across long spans of text, regardless of whether a visual cue is immediately available.
-
** lambda (The Hyperparameter):** lambda should be dynamically scheduled during training—starting low and increasing exponentially—to gradually increase the reliance on structural linguistic scaffolding over raw visual cues.
What the Improved AI System Can Do:
This system will achieve superior abstract reasoning and complex instruction following. By explicitly optimizing for long-range linguistic dependencies, the drafter moves beyond simple captioning or direct visual query answering. It can robustly handle multi-step reasoning problems, counterfactual questions (If X were changed, what would happen?
), and tasks requiring synthesis of information across disparate conceptual domains (e.g., combining knowledge from scientific papers with diagrams).
When processing a token t i:
-
Text-Only Mode: The module masks out all visual tokens, forcing the attention mechanism to rely purely on textual context and structural grammar derived from the text dataset.
-
Multimodal Mode: The module uses a learned confidence score based on the input prompt's complexity (e.g., high confidence for
Why is this?
vs. low confidence forWhat color is this?
). If the prompt complexity is high, SCMAM increases attention weight distribution towards text-derived tokens to stabilize reasoning and prevent over-reliance on potentially ambiguous visual details. -
Mixed Training: During mixed training, SCMAM ensures that the attention weights are structurally guided by the textual context first, and then refined by the visual input only when necessary for grounding.
-
Phase 1 (Foundation): Train exclusively on clean, high-quality Text-Only Data. This phase maximizes the acquisition of rich vocabulary and long-range dependency modeling (L dep optimization).
-
Phase 2 (Grounding): Introduce simple, direct Multimodal Tasks (e.g.,
Identify X in this image
). The objective here is to establish basic cross-modal alignment while preserving the linguistic structure learned in Phase 1. -
Phase 3 (Synthesis/Advanced): Introduce the full Mixed Data. The loss function weights must now balance the high structural demands of text with the grounding requirements of images. We should prioritize tasks requiring reasoning over visual evidence (e.g.,
Based on the principles described in this paragraph, what is likely to happen next?
).
Sources
- GPT-4 Technical Report
- Qwen2.5-VL Technical Report
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
- Accelerating Large Language Model Decoding with Speculative Sampling
- Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
- DREAM: Drafting with Refined Target Features and Entropy-Adaptive Cross-Attention Fusion for Multimodal Speculative Decoding
- Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- Speculative Decoding Reimagined for Multimodal Large Language Models
- ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
- TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
- Qwen2 Technical Report
- LLaMA: Open and Efficient Foundation Language Models
- VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Emu3: Next-Token Prediction is All You Need
- Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling
- MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
- Learning Harmonized Representations for Speculative Sampling
- Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks