Query Independent Variable Rate Visual Token Coding
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Query Independent Variable Rate Visual Token Coding".
Tom: Visual-token compression for vision–language models is posed almost entirely as a selection problem, but this work introduces a method to keep every token and vary its rate,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, let's talk about the specific title and authors of this piece, "Query Independent Variable Rate Visual Token Coding," because understanding who did this is important for where this research fits into the broader landscape of multimodal AI. The authors are Hongbo Zhang, Zihao Yang, Liuyang Song, Daqian Yang, Haoyang Yao, and Yan Wen.
Jane: I see the names here; it's a strong team coming from Peking University and Uppsala University and UC Santa Cruz. They clearly have a solid foundation in both the theoretical side of AI and the practical application side of vision systems.
Lu: Their focus seems to be on moving beyond simple selection problems, which is where most of the existing work has been stuck, trying to figure out which tokens are important based on a single query.
Meng: I'm interested in their methodology because if they've managed to keep every token and assign variable rates based on measured curves, that suggests a much more granular control over the compressed data than standard methods allow.
Lalam: That variable rate aspect is what makes it so appealing for deployment; it means the model isn't forced into a single compression level for everything, which could lead to better quality across different use cases.
The paper's summary: Tom: To summarize the main point of this paper, "Query Independent Variable Rate Visual Token Coding," they propose a method where you keep every token and adjust its rate independently based on its measured distortion-rate curve. This is significant because it solves the allocation problem exactly using dynamic programming to distribute a fixed bit budget across those curves.
Jane: That sounds like they took the selection problem—which is usually just deciding which tokens to drop—and transformed it into an allocation problem where you manage how much information each token gets within a set budget.
Lu: The authors are essentially creating a transform code that measures these distortion-rate curves for every token, and then the allocator uses dynamic programming to find the optimal rate index for each token that minimizes total distortion while respecting the fixed bit budget.
Meng: So it sounds like they've built a two-part system: one part to measure what each token is like, and another part to spend bits optimally on those measurements. That separation of concerns is interesting for practical engineering.
Lalam: And the implication I see immediately from that summary is that this resulting compressed representation can serve any future query without needing a new encoding step, which opens up possibilities for caching and transmission scenarios we’ve been talking about.
The paper's improvements: Tom: Now let's get into the actual improvements they detail; the paper focuses on how this approach handles allocation better than existing methods. They show that their method uses measured distortion-rate curves, not some parametric model of them, which is a key design decision because it fits the observed data better.
Jane: So instead of assuming a mathematical decay for how much distortion changes with rate, they are using empirical measurements to define those curves directly. That gives the allocation mechanism a much more grounded basis for making decisions.
Lu: They explicitly compare their exact dynamic programming solution to baselines like uniform rate coding and pruning methods, showing that their method spreads the budget across sixteen different levels, whereas pruning only uses two or uniform rate uses one level.
Meng: That difference in granularity is telling; if you can distribute the bits across sixteen distinct possibilities instead of just two choices, it suggests a much finer tuning of fidelity for different parts of the representation.
Lalam: And they show that this exact solution minimizes total distortion subject to a fixed bit budget by finding the optimal rate index for each token, which is a much more sophisticated way to use the allocated bits than simply setting uniform rates or using pruning scores.
Conclusion: Tom: So, wrapping things up on "Query Independent Variable Rate Visual Token Coding," the main takeaway is that by measuring distortion-rate curves and using exact dynamic programming for allocation, they get a representation that works regardless of the question you ask later. The performance gains are substantial when compared to uniform rate coding, showing reductions in output KL divergence of thirty-seven–forty-seven percent at the 4B model and 8B model.
Jane: It’s really impressive that they managed to achieve these fidelity improvements across multiple backbones and rates while keeping the allocation mechanism completely independent of the specific query being posed. This robustness is what makes this method so compelling for real-world deployment.
Lu: The transfer protocol they introduce, where you carry two independent human questions with an image, allows them to test exactly how much conditioning on a query costs once the representation is already out there, which gives a lot of insight into reusability.
Meng: From an engineering viewpoint, the fifty-eight ms average time it takes for that exact solve on the encoder side is totally manageable for practical use within an inference pipeline, which makes this feasible in a production environment.
Lalam: I think what really excites me about this paper is how this variable rate setting preserves fidelity even when you have to serve a completely different question later, which is crucial for any system that needs to be cached or transmitted across sessions.
Tom: That’s the essence of it; they’ve moved the allocation problem from a guessing game based on assumptions to an exact optimization based on measured data. So, we're leaving this paper with a method that offers superior performance and flexibility for vision-language models. Next up, we have some deep dives into how other teams are tackling ambiguity in multimodal translation.
Hongbo Zhang, *, Zihao Yang, *
Peking University · Uppsala University
cs.CV
Submitted: 2026-09-20
Updated: 2026-09-20
Importance score: 88/100
The gist: Visual-token compression for vision–language models is posed almost entirely as a selection problem, but this work introduces a method to keep every token and vary its rate, which solves the
Key concepts
- Transform Code
- This component measures each visual token's unique distortion-rate curve using PCA truncation error independence. This ensures the measured error accurately reflects how distortion changes with rate, regardless of the specific token being analyzed.
- Allocator
- The allocator uses exact dynamic programming to distribute a fixed bit budget across all tokens. It finds the optimal rate index for each token to minimize total distortion while respecting the overall bit limit, based on the measured curves.
- Query Independence
- RDTok's representation is query-independent because it measures token properties based only on the image itself, not a specific question. This means one compressed version can serve multiple different queries about the same image without needing to re-rank tokens.
- Measured Distortion-Rate Curve (Di(l))
- This curve quantifies how much distortion occurs for a specific token when it is assigned a particular bit rate (l). It is measured directly from the data, unlike parametric models that use assumptions about how distortion scales with rate.
Terminology
Summary
Visual-token compression for vision–language models is posed almost entirely as a selection problem, but this work introduces a method to keep every token and vary its rate, which solves the allocation problem exactly based on measured distortion–rate curves. This approach is significant because it creates a query-independent compressed representation that outperforms existing methods when the representation must serve different queries about the same image.
The gist
The resulting codec RDTok has two parts: a transform code that turns each token into a measured distortion–rate curve, and an allocator that spends a fixed bit budget across those curves by exact dynamic programming.
How it works
-
The transform code measures each token’s distortion–rate curve, denoted as Di(l), using PCA truncation error independence to ensure the error is independent of the rate.
-
For tokens from images disjoint from every evaluation image, a mean µ and an orthonormal basis U are fitted to code each token in the coefficient domain: yi = U(zi − µ), zˆi = U⊤yˆi + µ.
-
The distortion-rate curve Di(l) is then calculated for a rate index lρ using the formula: Di(l) = (zi − zˆi(lρ))2.
-
The allocation problem is solved exactly using dynamic programming to find the optimal rate index for each token, minimizing total distortion subject to a fixed bit budget B, as defined by Eq. (5): Jk(b) = min l∈L, l≤b n Jk−1(b − l) + Dk(l).
Allocation and Comparison
The allocator maps the measured curves D and the fixed budget B to a set of rate indices L, such that P i(l i) = B. The optimal allocation is found by minimizing XN i=1 Di(li) s.t. X i li = B (Eq. 4). This exact solution is computationally tractable, taking only O(NBL) time, which is on the encoder-side only and takes approximately 58 ms on average for a single solve.
Performance Gains
The RDTok method demonstrates superior performance across multiple metrics compared to existing baselines when serving different queries about the same image. Across three rates (256/384/512 bits per token) and two backbones (4B and 8B Qwen3-VL-Instruct), RDTok consistently lowers output KL in every cell and leads on other quantities in all but a couple of COCO cells, with margins widening as compression becomes more aggressive. For instance, against uniform rate coding, RDTok lowers output KL by 37–47% at 4B and 43–45% at 8B.
Query Independence
The key advantage is that the measured curves D are properties of the image alone; they are not conditioned on the query. Unlike pruning methods, which rank tokens based on a question, or attention-ranked pruning, which reads the query to select tokens about as well as rate allocation spends bits, RDTok's allocation is query-independent by construction.
This means that one compressed representation serves any subsequent query without needing to re-evaluate the ranking. The advantage becomes decisive when the compressed image must answer a different question about the same image, where attention-ranked pruning loses fidelity.
Analysis of Allocation
The analysis shows that RDTok spreads the budget over sixteen levels, whereas pruning has two by construction and uniform rate one. The closed-form water-fill collapses onto three levels because it assumes a parametric decay that does not match the measured curves. The gain is derived from taking distortion–rate curves from measurement rather than from a parametric model, which is the central design decision of the paper. Furthermore, weighting Eq. (4) by each region’s measured effect on the output KL allows for an allocation that measures something real—a property of the query it came from—but this weighting does not transfer to the next question about the same image. The source-MSE allocator is shown to be optimizing a surrogate, and spreading error over multiple tokens is superior to concentrating it on a quarter of them.
Transfer Protocol
The paper introduces a transfer protocol by carrying two independent human questions with every ChartQA test image, allowing for the measurement of what conditioning an allocation on the query costs once the representation outlives that query. This asymmetry—where one property depends on the query and another is fixed by the image—is what grants RDTok its advantage in cached or transmitted scenarios. The larger model’s attention is more query-specific, meaning a larger share of what pruning selects is particular to the question asked, making reusability precisely what this variable-rate setting preserves.
Improvements for AI systems
Here are the specific improvements to AI systems achievable by implementing the methods described in this paper, along with what those improved systems can do:
-
The development of a new visual-token codec, called RDTok (Rate-Distortion Tokenizer), which is query-independent and variable-rate.
-
Implementation of an exact bit allocation mechanism using dynamic programming to distribute a fixed bit budget across tokens based on their measured distortion–rate curves, rather than relying on binary selection or uniform rate assignment.
-
A transfer protocol that allows a single compressed visual representation to serve multiple, potentially different queries without needing to re-encode or re-rank based on the new question.
The improved AI system (Vision-Language Model) can perform the following specific tasks:
-
It can maintain a single, highly compressed visual representation of an image that is written once and read many times (e.g., cached across conversation turns or stored for corpus querying).
-
When presented with a new question, the system can dynamically allocate bits to each visual token based on how much distortion (loss of information) is incurred versus the rate/cost saved, optimizing this allocation precisely for the specific query being asked, without needing to know what that future question will be in advance.
-
It achieves superior performance across various compression levels compared to traditional methods like uniform-rate coding and token pruning by spending bits intelligently based on measured image properties rather than assumptions about the model's attention weights or parametric decay models.
-
The system demonstrates improved output preservation (higher accuracy) and lower output KL divergence (better fidelity to the original answer) across multiple benchmarks (ChartQA, COCO) at a fixed bit budget, achieving gains of 37–47% over state-of-the-art baselines like attention-ranked pruning.
-
It exhibits robust performance in scenarios where the compressed representation must serve a query completely different from the one used to generate it (query transfer), maintaining high fidelity where other methods degrade significantly (e.g., attention-ranked pruning loses over half its fidelity).
Sources
- LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
- RUTA: Principled Visual Token Allocation via Rate-Utility Optimization
- Adaptive-VoCo: Complexity-Aware Visual Token Compression for Vision-Language Models
- OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning
- Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models
- Qwen3-VL Technical Report
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models