LensVLM: Selective Context Expansion for Compressed Visual Representation of Text

arXiv:2605.07019 · cs.CV, cs.AI · Submitted 2026-05-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "LensVLM: Selective Context Expansion for Compressed Visual Representation of Text".

Jane: LensVLM introduces an inference framework and post-training recipe that enables Vision Language Models (VLMs) to scan compressed images and selectively expand only the relevant images to their uncompressed form via…

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, to recap what we’ve heard, we’re talking about "LensVLM: Selective Context Expansion for Compressed Visual Representation of Text," which is basically an inference framework designed to help VLMs handle text as compressed images. The main thesis here is that while varying rendering resolution lets you compress the text representation, accuracy usually drops quickly because characters get lost at higher compression rates.

Jane: Exactly, and the authors propose LensVLM as a solution by enabling the model to first scan these compressed images and then use a learned tool to selectively expand only the relevant parts of those images back into their uncompressed form when necessary. This is different from older methods that decide what to keep before they start; this approach makes the expansion decision on demand during inference.

Lu: It’s a really clever idea because it acknowledges the limitation that traditional compression methods have—they lose information irreversibly once a certain threshold is crossed, and LensVLM builds a way around that loss by only spending resources on the parts of the input that matter for the final answer.

Meng: So, essentially, they are giving the VLM an "Expand tool" it can call during its reasoning process to recover high-resolution text or images only when its internal representation indicates those specific details are needed. That’s a concrete mechanism we can actually build with a policy layer.

Lalam: That selective recovery mechanism is what really excites me because it suggests that future AI systems won't just be about processing everything at once; they’ll be about intelligent resource allocation for visual data, which could lead to much more efficient and targeted AI applications in the future.

Tom: And the paper highlights their training process too; they use a three-stage post-training recipe: supervised fine-tuning to teach it from ground truth traces, followed by reinforcement learning to optimize its policy based on maximizing correctness and tool usage. That’s how they instill this selective behavior.

Jane: Right, and the results they report are quite strong; LensVLM maintains accuracy comparable to the full-text upper bound even at a four point three times effective compression rate, which is way better than what previous retrieval-based or text and visual compression baselines could achieve up to ten point one times compression.

Lu: The paper notes that training makes the visual compression robust to rendering choices; they found that it reduced an eighteen-point accuracy spread across different rendering configurations down to under one point, which shows the model is learning a generalized way to handle those input variations.

Meng: That robustness is important for real-world deployment because it means we don't need perfect pre-processing of the input images; the system can handle slight variations in how the text was initially rendered without failing.

Lalam: It really shows that by training this way, we aren't just getting a better compression ratio; we’re getting a more resilient AI that can handle real-world messy inputs without losing critical context.

Conclusion: Tom: So, wrapping up our discussion on "LensVLM: Selective Context Expansion for Compressed Visual Representation of Text," the title itself really captures the essence of what they achieved: selective context expansion for compressed visual representation of text. The authors, Roy Xie and his team, developed a framework that allows VLMs to intelligently manage information access during inference.

Jane: It’s about moving away from a model that blindly processes everything at once and instead giving it a tool to decide precisely which parts of the input need to be fully expanded for an accurate response. In simple terms, it means the AI learns when to spend its processing power on detail versus when it can rely on the compressed representation.

Lu: The implication is that we are moving toward more sophisticated AI systems that don't just process data passively; they actively manage their cognitive load based on what’s needed for the task, which opens up avenues for much deeper levels of contextual understanding.

Meng: From a practical standpoint, this suggests that future AI applications will need to incorporate intelligent resource management into their core design, ensuring the system knows how to dynamically trade off visual detail against computational cost effectively.

Lalam: And for the broader impact, I think this points toward a future where AI can interact with massive amounts of visual information in ways that are both highly efficient and extremely accurate because it’s not trying to process every pixel all the time.

Tom: That’s a big picture idea, Jane; essentially, we are seeing AI develop strategies for intelligent information filtering at the point of use rather than relying on brute force processing of everything. It seems like this work lays a solid foundation for building more nuanced and resource-aware visual reasoning systems.

Apple

cs.CV, cs.AI

Submitted: 2026-05-07

Updated: 2026-10-02

Comments: Accepted to NeurIPS 2026

Code: https://github.com/apple/ml-lensvlm

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 92/100

The gist: LensVLM introduces an inference framework and post-training recipe that enables Vision Language Models (VLMs) to scan compressed images and selectively expand only the relevant images to their

Key concepts

Deterministic Renderer
This component converts raw text into a sequence of images based on specific settings like font and spacing. It acts as a controlled way to simulate different levels of image compression, allowing the model to test how much detail it can still extract from the visual representation.
Expand Tool
A learned tool that the VLM uses during inference to recover selected images from their compressed state into their full, uncompressed form. The model learns when and where to use this tool based on its reasoning needs, such as needing more detail for a specific part of the text.
Information-Theoretic Tradeoff
A mathematical analysis quantifying the balance between reading compressed visual data directly and using the Expand tool. It predicts that tool use is most beneficial when direct reading accuracy drops faster than selection accuracy, defining an optimal compression regime for performance.

Terminology

Summary

LensVLM introduces an inference framework and post-training recipe that enables Vision Language Models (VLMs) to scan compressed images and selectively expand only the relevant images to their uncompressed form via learned tools, pushing compression further than baselines while maintaining high accuracy.

How it works

The core idea is to treat text as an image that can be compressed by varying rendering resolution, which allows for fine-grained compression knobs. LensVLM addresses the issue where accuracy deteriorates quickly as compression increases because characters shrink below the vision encoder’s effective resolution, making them indistinguishable. The framework operates in two main stages: first, the model scans compressed images, and second, it uses a learned tool to selectively expand only the relevant images to their uncompressed form via learned tools. This approach is distinct from retrieval methods that preprocess and commit to what is retained offline; LensVLM takes compressed visual tokens as input and expands only the images the model selects on demand.

Key Components and Training Pipeline

The LensVLM process involves several key components:

  1. A deterministic renderer, parameterized by a rendering configuration (font, geometry, line spacing), converts text into a sequence of images.

  2. A vision encoder maps each image to visual tokens; the total input compression rate is defined as the ratio of text tokens to visual tokens consumed (ICR).

  3. The VLM policy samples a multi-turn trajectory where it interleaves reasoning with an Expand tool call, which recovers selected compressed images into uncompressed form (source text or high-resolution images).

To teach this behavior, LensVLM employs a three-stage post-training recipe:

  1. Supervised Fine-Tuning (SFT): The model is trained on synthetic tool-use traces generated with ground-truth evidence, minimizing the autoregressive loss over the produced tokens.

  2. Reinforcement Learning (RL): Starting from the SFT checkpoint, on-policy RL optimizes the policy to maximize an expected reward function that incentivizes correct answer and tool-use behavior. The reward is defined as:

R(y, a⋆) = 0.7 · c + 0.3 · c · u, where 'c' is answer correctness and 'u' indicates whether the trajectory uses the Expand tool.

Empirical Findings on Performance

LensVLM demonstrates significant performance gains across various benchmarks:

(Figure 1 shows that LensVLM maintains accuracy comparable to the full-text upper bound at 4.3× effective compression and outperforms retrieval-based, text- and visual-compression baselines up to 10.1× effective compression across seven text QA benchmarks.)

The analysis validates the approach by showing that training makes visual compression robust to rendering choices, reducing an 18-point accuracy spread across rendering configurations to under 1 point. Furthermore, attention analysis shows that training redirects attention away from distractor images and toward the content returned by Expand, with this shift growing at higher compression. The model learns that text expansion is preferable for rendered text, while high-resolution image expansion suits native documents whose layout cues carry task-relevant information.

Tool Modality and Generalization

The study investigates the modality of the expanded content. Findings show that the bottleneck is information access, not reasoning: chain-of-thought without tool access provides no benefit over direct answers, as reasoning over illegible images might introduce noise. The best tool response at every compression level is original text, which outperforms OCR and image zoom on rendered text. However, the ranking reverses on native multimodal documents: the zoom tool outperforms text tool because text tool discards layout and visual structure that the high-resolution image preserves.

LensVLM also generalizes beyond natural-language text. On code understanding benchmarks (RepoQA, CodeQueries), it improves from Comp. Image and outperforms Glyph at every compression level, demonstrating that tool-use transfers to code without code-specific training. This suggests the framework is a general-purpose inference strategy applicable to frontier VLMs through prompting alone.

Efficiency and Information Theory

The paper analyzes the trade-off between visual reading and tool use from an information-theoretic perspective. The benefit of selective expansion is quantified by the formula: Dno(ρ) − D¯(π) = pπ [Dno(ρ) − Dhit] − (1 − pπ) [Dmiss − Dno(ρ)], which makes the tradeoff explicit. The analysis predicts a useful compression regime: tool use helps when direct reading degrades faster than selection accuracy, a pattern confirmed across all studied compression rates. Efficiency analysis shows that while LensVLM incurs overhead due to multi-turn sequential decoding and vision encoder processing, it achieves significant token reduction compared to text-only baselines, with the final turn being the peak KV cache occupancy.

Improvements for AI systems

Here are specific improvements for AI systems based on the LensVLM framework:

  1. Improve long-context reasoning accuracy in document QA tasks by enabling models to selectively expand only relevant visual evidence instead of relying on fixed, compressed views or full-text retrieval.

  2. Enhance robustness against varying rendering configurations by training VLMs to be invariant to font/layout choices, allowing the model to rely on the content of expanded text rather than visual legibility when compression is high.

  3. Develop multimodal document understanding systems that can effectively utilize native document layouts by employing a zoom tool to preserve layout cues, outperforming OCR-based text extraction when visual structure is crucial for answering questions about spatial relationships or document context.

  4. Create zero-shot transfer capabilities for code understanding and retrieval tasks by teaching VLMs to use image expansion tools on rendered code blocks without needing task-specific training data.

  5. Implement a tool-choice mechanism that intelligently selects between expanding text (for rendered text) and zooming (for native documents), optimizing the retrieval of evidence based on the specific visual context of the document chunk being processed.

  6. Create an inference framework that balances accuracy recovery with efficiency by using a learned selection mechanism to identify relevant images from compressed inputs, minimizing unnecessary expansion calls while maximizing gains in complex multi-hop reasoning.

  7. Improve model scalability by identifying the minimum required parameter size (e.g., 9B) necessary for reliable agentic tool-use behavior, ensuring that smaller models are appropriately augmented with tool capabilities to achieve high performance.

Abstract

Vision Language Models (VLMs) offer the exciting possibility of processing text as rendered images, bypassing the need for tokenizing the text into long token sequences. Since VLM image encoders map fixed-size images to a fixed number of visual tokens, varying rendering resolution provides a fine-grained compression knob. However, accuracy deteriorates quickly as compression increases: characters shrink below the vision encoder's effective resolution, making them indistinguishable. To address this, we propose LensVLM, an inference framework and post-training recipe that enables VLMs to scan compressed images, then selectively expand only the relevant images to their uncompressed form via learned tools. Building on Qwen3.5-9B-Base, LensVLM maintains accuracy comparable to the full-text upper bound at 4.3 times effective compression and outperforms retrieval-based, text- and visual-compression baselines up to 10.1 times effective compression across seven text QA benchmarks. LensVLM also generalizes to multimodal document and code understanding tasks, with the accuracy gain over baselines growing as compression increases. Our analysis validates this approach: training makes visual compression robust to rendering choices, and as compression grows the model increasingly relies on expanded content rather than unreliable visual reading. The analysis also yields practical tool-choice guidance: text expansion is preferable for rendered text, while high-resolution image expansion suits native documents whose layout cues carry task-relevant information.

Sources

Related papers